<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>LatentEval</title><description>Independent research on AI agent reliability: how multi-agent failures cascade and propagate, and whether the evals meant to certify them hold up statistically.</description><link>https://latenteval.ai</link><language>en</language><atom:link href="https://latenteval.ai/rss.xml" rel="self" type="application/rss+xml"/><item><title>Agentic RAG architecture, and where each pattern breaks</title><link>https://latenteval.ai/analysis/agentic-rag-architecture</link><guid isPermaLink="true">https://latenteval.ai/analysis/agentic-rag-architecture</guid><description>Agentic RAG wraps retrieval in an autonomous agent that plans, routes, and self-corrects. Here is the architecture, the failure modes each pattern adds, and when classic RAG is the safer call.</description><pubDate>Fri, 10 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Every tool behind one MCP gateway. One breach reaches all of them.</title><link>https://latenteval.ai/building-ai/mcp-gateway</link><guid isPermaLink="true">https://latenteval.ai/building-ai/mcp-gateway</guid><description>An MCP gateway fronts many MCP servers as one endpoint, brokering credentials and policy. It buys one audit trail and concentrates every tool into one process. Gateway vs proxy, when to run one.</description><pubDate>Fri, 10 Jul 2026 00:00:00 GMT</pubDate></item><item><title>The agent reliability testing checklist your green eval skips</title><link>https://latenteval.ai/analysis/agent-reliability-testing-checklist</link><guid isPermaLink="true">https://latenteval.ai/analysis/agent-reliability-testing-checklist</guid><description>A pre-flight agent reliability testing checklist: six checks with pass criteria and calculators, from failure-mode coverage and statistical power to pass^k, cascade budget, and judge calibration.</description><pubDate>Fri, 10 Jul 2026 00:00:00 GMT</pubDate></item><item><title>The best Claude model for coding is rarely the one at the top.</title><link>https://latenteval.ai/building-ai/which-claude-model</link><guid isPermaLink="true">https://latenteval.ai/building-ai/which-claude-model</guid><description>Which Claude model should you use? A task-by-task routing guide across Opus 4.8, Sonnet 5, Haiku 4.5, and Fable 5, with a checkability rule for coding, writing, research, and agents.</description><pubDate>Fri, 10 Jul 2026 00:00:00 GMT</pubDate></item><item><title>The LLM-judge bias checklist that gates your ranking</title><link>https://latenteval.ai/analysis/llm-as-a-judge-bias-checklist</link><guid isPermaLink="true">https://latenteval.ai/analysis/llm-as-a-judge-bias-checklist</guid><description>The LLM-judge bias checklist is eight pass/fail gates you run before trusting a ranking. Each gate pairs a detection test with a numeric pass line and the calculator that computes it.</description><pubDate>Fri, 10 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Langfuse vs LangSmith: the real split, and what all three still miss</title><link>https://latenteval.ai/building-ai/langfuse-vs-langsmith</link><guid isPermaLink="true">https://latenteval.ai/building-ai/langfuse-vs-langsmith</guid><description>A verified comparison of Langfuse, LangSmith, and Braintrust on licensing, self-hosting, tracing, evals, and pricing shape, plus the statistical eval layer none of the three computes for you.</description><pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Opus vs Sonnet: the routing question your task already answered</title><link>https://latenteval.ai/building-ai/opus-vs-sonnet</link><guid isPermaLink="true">https://latenteval.ai/building-ai/opus-vs-sonnet</guid><description>Claude Opus 4.8 lists at 1.67x Sonnet 5, yet on checkable work our routing eval measured no separation between them. A table for routing Opus, Sonnet, Haiku, and Fable 5 by task risk and cost.</description><pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Model routing and the refusal tax: a pre-registered study</title><link>https://latenteval.ai/research/model-routing-refusal-tax</link><guid isPermaLink="true">https://latenteval.ai/research/model-routing-refusal-tax</guid><description>On short, checkable tasks we measured no Opus 4.8-to-Fable 5 capability separation at either effort we ran, and the premium tier’s refusal ‘rescue’ silently served the cheaper model on 20 of 28 calls.</description><pubDate>Tue, 07 Jul 2026 00:00:00 GMT</pubDate></item><item><title>OpenAI Evals is winding down. The alternatives skip the statistics.</title><link>https://latenteval.ai/building-ai/openai-evals</link><guid isPermaLink="true">https://latenteval.ai/building-ai/openai-evals</guid><description>OpenAI is deprecating its hosted Evals platform and steering users to Promptfoo, which it is acquiring. What OpenAI Evals, DeepEval, Ragas, TruLens, and Promptfoo each do, and the rigor none ship.</description><pubDate>Sat, 04 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Amazon v. Perplexity: the ruling that blocked an agent you authorized</title><link>https://latenteval.ai/using-ai/amazon-perplexity-agent-ruling</link><guid isPermaLink="true">https://latenteval.ai/using-ai/amazon-perplexity-agent-ruling</guid><description>You told your shopping agent yes, and Amazon wants a court to say no; it won that order against Perplexity&apos;s Comet, and the Ninth Circuit has not yet ruled on the appeal.</description><pubDate>Fri, 03 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Do you need Fable 5? A task-by-task verdict, every number sourced</title><link>https://latenteval.ai/using-ai/do-you-need-fable-5</link><guid isPermaLink="true">https://latenteval.ai/using-ai/do-you-need-fable-5</guid><description>Claude Fable 5 is included on Pro, Max, and Team through July 7, then metered at twice Opus 4.8&apos;s price. Which tasks justify switching it on, and which are better left on Opus 4.8.</description><pubDate>Fri, 03 Jul 2026 00:00:00 GMT</pubDate></item><item><title>SEP-2567 deletes Mcp-Session-Id, and sticky routing with it</title><link>https://latenteval.ai/building-ai/mcp-goes-stateless</link><guid isPermaLink="true">https://latenteval.ai/building-ai/mcp-goes-stateless</guid><description>Roots, Sampling, and Logging are only annotation-deprecated; SEP-2567 deletes Mcp-Session-Id and sticky routing on July 28 with no grace period. Audit the session layer first.</description><pubDate>Fri, 03 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Model fallback swaps the model unseen, and refusals log as success</title><link>https://latenteval.ai/building-ai/log-refusals-and-model-fallbacks</link><guid isPermaLink="true">https://latenteval.ai/building-ai/log-refusals-and-model-fallbacks</guid><description>When Claude Fable 5 refuses, it returns HTTP 200 with an empty content array, so error dashboards read it as a success, and with fallback on the model can swap to Opus 4.8 unseen. Instrument now.</description><pubDate>Thu, 02 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Which Claude model for your task, and what reaching up really costs</title><link>https://latenteval.ai/building-ai/routing-claude-models-by-task</link><guid isPermaLink="true">https://latenteval.ai/building-ai/routing-claude-models-by-task</guid><description>On short work you can check, our routing eval found no Opus-to-Fable capability separation at both effort levels. Buy down: Fable 5&apos;s 2x buys a refusal tax and a fallback to the cheaper model.</description><pubDate>Thu, 02 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Make AI agents reliable in production by budgeting for failure</title><link>https://latenteval.ai/analysis/reliable-ai-agents-in-production</link><guid isPermaLink="true">https://latenteval.ai/analysis/reliable-ai-agents-in-production</guid><description>Reliable AI agents in production start with the compounding math: 0.95 per step over 20 steps is ~36% end to end. A topology-aware playbook of levers: retry, fallback, checkpoint, human-gate, degrade.</description><pubDate>Tue, 30 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Where RAGAS wins at RAG evaluation, and where it stops</title><link>https://latenteval.ai/analysis/ragas-rigorous-rag-evaluation</link><guid isPermaLink="true">https://latenteval.ai/analysis/ragas-rigorous-rag-evaluation</guid><description>A positioning read on RAGAS for RAG evaluation: its real strengths, a limitation its own authors flag, and the rigor gaps around confidence intervals, significance, and propagation-aware attribution.</description><pubDate>Tue, 30 Jun 2026 00:00:00 GMT</pubDate></item><item><title>The OWASP AI agent security cheat sheet, read as a reliability spec</title><link>https://latenteval.ai/building-ai/owasp-ai-agent-security-cheat-sheet</link><guid isPermaLink="true">https://latenteval.ai/building-ai/owasp-ai-agent-security-cheat-sheet</guid><description>OWASP&apos;s AI Agent Security Cheat Sheet and the 2026 Agentic Top 10 name the same agent failures. Each maps to a reliability control you can test: least-privilege tools, isolated memory, bounded loops.</description><pubDate>Tue, 23 Jun 2026 00:00:00 GMT</pubDate></item><item><title>How to evaluate a RAG pipeline beyond a single score</title><link>https://latenteval.ai/analysis/how-to-evaluate-rag-pipeline</link><guid isPermaLink="true">https://latenteval.ai/analysis/how-to-evaluate-rag-pipeline</guid><description>Evaluate a RAG pipeline by scoring retrieval and generation separately, putting a confidence interval on every metric, and attributing each failure to the stage that produced it.</description><pubDate>Mon, 22 Jun 2026 00:00:00 GMT</pubDate></item><item><title>RAG pipeline failure modes, and the gate that stops each one</title><link>https://latenteval.ai/analysis/rag-pipeline-failure-modes</link><guid isPermaLink="true">https://latenteval.ai/analysis/rag-pipeline-failure-modes</guid><description>RAG pipeline failure modes are containment failures at the retrieval-to-generation boundary: five modes mapped to how each propagates, its detection signal, and the gate that holds it.</description><pubDate>Mon, 22 Jun 2026 00:00:00 GMT</pubDate></item><item><title>AI observability proves the run finished. An eval proves it was right.</title><link>https://latenteval.ai/building-ai/observability-is-not-evals</link><guid isPermaLink="true">https://latenteval.ai/building-ai/observability-is-not-evals</guid><description>LangChain&apos;s State of Agent Engineering survey of 1,340 practitioners: 89% run observability, only 52.4% run offline evals. An offline eval is what scores whether the answer was right.</description><pubDate>Wed, 17 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Do microservices resilience patterns port to AI agents?</title><link>https://latenteval.ai/analysis/cascading-failure-microservices-to-agents</link><guid isPermaLink="true">https://latenteval.ai/analysis/cascading-failure-microservices-to-agents</guid><description>Five proven microservices resilience patterns, from circuit breaker to timeout budget, mapped to their AI agent equivalents, with a judgment on how far each analogy actually holds.</description><pubDate>Wed, 17 Jun 2026 00:00:00 GMT</pubDate></item><item><title>What your AI remembers, and whether it trains on your chats</title><link>https://latenteval.ai/using-ai/what-your-ai-remembers</link><guid isPermaLink="true">https://latenteval.ai/using-ai/what-your-ai-remembers</guid><description>You flipped the switch you could see. The other one is still on.</description><pubDate>Wed, 17 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Cascading failures in agent systems, from trigger to containment</title><link>https://latenteval.ai/analysis/cascading-failures-agent-systems</link><guid isPermaLink="true">https://latenteval.ai/analysis/cascading-failures-agent-systems</guid><description>OWASP&apos;s ASI08 files cascading failures under security. Reframed as error-propagation engineering, one fault becomes two measurable quantities, propagation radius and containment rate, per topology.</description><pubDate>Wed, 10 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Multi-agent orchestration patterns and the failures they amplify</title><link>https://latenteval.ai/analysis/multi-agent-orchestration-patterns</link><guid isPermaLink="true">https://latenteval.ai/analysis/multi-agent-orchestration-patterns</guid><description>The five multi-agent orchestration patterns, supervisor, sequential-pipeline, swarm, debate, and blackboard, mapped to how errors cascade in each and the failure modes each one amplifies.</description><pubDate>Wed, 10 Jun 2026 00:00:00 GMT</pubDate></item><item><title>The prompt wording is a hyperparameter you never swept.</title><link>https://latenteval.ai/building-ai/the-lazy-prompter-problem</link><guid isPermaLink="true">https://latenteval.ai/building-ai/the-lazy-prompter-problem</guid><description>Rewording the same task swings a model&apos;s pass rate: format, option order, even a &apos;please&apos;. A one-phrasing eval samples one point from a spread you never measured. Pin the prompt and measure it.</description><pubDate>Wed, 10 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Agentic AI testing beyond a single eval run</title><link>https://latenteval.ai/analysis/agent-reliability-testing</link><guid isPermaLink="true">https://latenteval.ai/analysis/agent-reliability-testing</guid><description>Single-run eval samples agent reliability once. Rigorous testing measures it across many runs with confidence intervals, statistical power, pass^k, and fault injection for cascade propagation.</description><pubDate>Wed, 03 Jun 2026 00:00:00 GMT</pubDate></item><item><title>How to measure agent reliability past a single pass rate</title><link>https://latenteval.ai/analysis/how-to-measure-agent-reliability</link><guid isPermaLink="true">https://latenteval.ai/analysis/how-to-measure-agent-reliability</guid><description>How to measure agent reliability with metrics that capture the consistency a single pass rate cannot: pass@k versus pass^k, a reliability@k suite aggregate, and a confidence interval on every rate.</description><pubDate>Wed, 03 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Is your eval difference statistically significant?</title><link>https://latenteval.ai/analysis/is-my-eval-statistically-significant</link><guid isPermaLink="true">https://latenteval.ai/analysis/is-my-eval-statistically-significant</guid><description>Two eval runs a few points apart. Separate a real gain from run-to-run noise with a paired McNemar test on the same items: a p-value and a confidence interval on the pass-rate delta.</description><pubDate>Wed, 03 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Is Chrome&apos;s agentic browser safe when it clicks Buy for you?</title><link>https://latenteval.ai/using-ai/chrome-ai-browsing-safety</link><guid isPermaLink="true">https://latenteval.ai/using-ai/chrome-ai-browsing-safety</guid><description>Google&apos;s auto browse lets an AI click and type across your tabs, in a US-only paid preview. The same agent obeys instructions hidden on the pages it reads. How to keep a hand on it.</description><pubDate>Wed, 27 May 2026 00:00:00 GMT</pubDate></item><item><title>Whose side is your AI shopping agent on?</title><link>https://latenteval.ai/using-ai/whose-side-is-your-shopping-agent-on</link><guid isPermaLink="true">https://latenteval.ai/using-ai/whose-side-is-your-shopping-agent-on</guid><description>Your AI shopping agent has your card and picks what you buy. Google, Amazon, and Perplexity all ship one now, and when a model carries a sponsor&apos;s incentive, it steers. Learn whose side yours is on.</description><pubDate>Wed, 27 May 2026 00:00:00 GMT</pubDate></item><item><title>How benchmarks get gamed, and how to check yours</title><link>https://latenteval.ai/building-ai/how-benchmarks-get-gamed</link><guid isPermaLink="true">https://latenteval.ai/building-ai/how-benchmarks-get-gamed</guid><description>BenchJack, a Berkeley auditing tool, found 219 flaws in ten popular agent benchmarks and gamed nine to near-perfect scores. Why a benchmark number is a claim about the harness, and how to check it.</description><pubDate>Sat, 16 May 2026 00:00:00 GMT</pubDate></item><item><title>How many runs a reliable eval needs to catch a regression</title><link>https://latenteval.ai/analysis/how-many-runs-for-a-reliable-eval</link><guid isPermaLink="true">https://latenteval.ai/analysis/how-many-runs-for-a-reliable-eval</guid><description>How many runs a reliable eval needs is a power calculation set by the regression you must catch, your target power, and the baseline pass rate. Includes a runs-needed table and the formula behind it.</description><pubDate>Sat, 16 May 2026 00:00:00 GMT</pubDate></item><item><title>Why AI sounds confident when it&apos;s wrong</title><link>https://latenteval.ai/using-ai/why-ai-sounds-confident-when-wrong</link><guid isPermaLink="true">https://latenteval.ai/using-ai/why-ai-sounds-confident-when-wrong</guid><description>A language model&apos;s confidence reads like clean handwriting: the page stays just as neat whether the claim underneath is solid or hollow. Why right and wrong arrive in the same voice.</description><pubDate>Sat, 16 May 2026 00:00:00 GMT</pubDate></item><item><title>AI agent evaluation that follows the whole trajectory</title><link>https://latenteval.ai/analysis/ai-agent-evaluation</link><guid isPermaLink="true">https://latenteval.ai/analysis/ai-agent-evaluation</guid><description>AI agent evaluation breaks when it scores the final answer and skips the path. Evaluate the trajectory, catch early-step corruption, and report pass rates with intervals.</description><pubDate>Sat, 09 May 2026 00:00:00 GMT</pubDate></item><item><title>Bias-correct your LLM-as-a-judge eval before reporting it</title><link>https://latenteval.ai/analysis/reporting-llm-as-a-judge-evaluations</link><guid isPermaLink="true">https://latenteval.ai/analysis/reporting-llm-as-a-judge-evaluations</guid><description>An LLM judge is an imperfect classifier, so its raw pass rate is biased. Correct it with the judge&apos;s sensitivity and specificity, then report a calibration-aware confidence interval.</description><pubDate>Sat, 09 May 2026 00:00:00 GMT</pubDate></item><item><title>AI agent reliability, from consistency to containment</title><link>https://latenteval.ai/analysis/ai-agent-reliability</link><guid isPermaLink="true">https://latenteval.ai/analysis/ai-agent-reliability</guid><description>AI agent reliability is a discipline of five properties: consistency, robustness, predictability, safety, and error propagation, with a map of where each is measured.</description><pubDate>Sat, 02 May 2026 00:00:00 GMT</pubDate></item><item><title>LLM-as-a-judge bias, and the tests that catch it</title><link>https://latenteval.ai/analysis/llm-as-a-judge-bias</link><guid isPermaLink="true">https://latenteval.ai/analysis/llm-as-a-judge-bias</guid><description>LLM-as-a-judge bias is systematic, measurable distortion in an evaluator. A per-bias map pairs each bias with a detection test and a correction, so you can tell when a judge&apos;s ranking would flip.</description><pubDate>Sat, 02 May 2026 00:00:00 GMT</pubDate></item><item><title>What LLM evals are, and what each type can certify</title><link>https://latenteval.ai/analysis/what-are-llm-evals</link><guid isPermaLink="true">https://latenteval.ai/analysis/what-are-llm-evals</guid><description>LLM eval covers four instruments: offline benchmark, LLM-as-judge, human, and online, each answering a different question, plus the benchmark-vs-product line and the rigor behind a trustworthy score.</description><pubDate>Sat, 02 May 2026 00:00:00 GMT</pubDate></item><item><title>Is your LLM-as-a-judge reliable? Test the evaluator</title><link>https://latenteval.ai/analysis/llm-as-a-judge</link><guid isPermaLink="true">https://latenteval.ai/analysis/llm-as-a-judge</guid><description>An LLM-as-a-judge is a fallible evaluator. Its reliability breaks along three axes, agreement, calibration, and bias, each with a test and a correction. This hub routes to all three.</description><pubDate>Sat, 18 Apr 2026 00:00:00 GMT</pubDate></item><item><title>Multi-agent LLM failure modes, and how to contain error propagation</title><link>https://latenteval.ai/analysis/multi-agent-failure-modes</link><guid isPermaLink="true">https://latenteval.ai/analysis/multi-agent-failure-modes</guid><description>Why multi-agent LLM systems fail, grounded in the MAST failure taxonomy and mapped to how each failure propagates across agent topologies and the containment levers that bound the propagation radius.</description><pubDate>Sat, 18 Apr 2026 00:00:00 GMT</pubDate></item><item><title>Prompt injection by calendar invite, and how to stop it</title><link>https://latenteval.ai/using-ai/calendar-invite-prompt-injection</link><guid isPermaLink="true">https://latenteval.ai/using-ai/calendar-invite-prompt-injection</guid><description>Your AI reads your calendar, and a stranger&apos;s invite can quietly hand it orders you never see.</description><pubDate>Sat, 18 Apr 2026 00:00:00 GMT</pubDate></item><item><title>LLM evals: which methods to trust and where they lie</title><link>https://latenteval.ai/analysis/llm-evals</link><guid isPermaLink="true">https://latenteval.ai/analysis/llm-evals</guid><description>LLM evals report whether a model passed. Whether that score is valid is a separate question. This hub maps the eval methods and the four ways an eval number lies, each routed to its fix.</description><pubDate>Tue, 07 Apr 2026 00:00:00 GMT</pubDate></item><item><title>Multi-agent systems, defined by how they fail</title><link>https://latenteval.ai/analysis/multi-agent-systems</link><guid isPermaLink="true">https://latenteval.ai/analysis/multi-agent-systems</guid><description>A multi-agent system is defined by its failure surface: agent, orchestration, coordination, shared state, and topology, each defined through the failure it enables, then routed to the research.</description><pubDate>Tue, 07 Apr 2026 00:00:00 GMT</pubDate></item></channel></rss>