LatentEval

FAILURE MODES | PROPAGATION | EVAL RIGOR

AI agent reliability, measured beyond the eval.

Independent research on how agentic systems fail: how errors cascade across agents, how far they propagate, and whether the evaluations meant to certify them hold up.

  • Research with confidence intervals
  • Cited analyses of how agents fail
  • Nineteen tools for eval statistics
In order

Essential reading

Five pieces in reading order. Start at the top if you are new here.

  1. Research

    Agentic AI testing beyond a single eval run

    Single-run eval samples agent reliability once. Rigorous testing measures it across many runs with confidence intervals, statistical power, pass^k, and fault injection for cascade propagation.

  2. Research

    LLM-as-a-judge bias, and the tests that catch it

    LLM-as-a-judge bias is systematic, measurable distortion in an evaluator. A per-bias map pairs each bias with a detection test and a correction, so you can tell when a judge's ranking would flip.

  3. Guides

    We scored Fable 5 vs Opus 5 vs Opus 4.8 and only the old one answered every call

    While Claude Fable 5 and Claude Opus 5 are tied on our reliability benchmarks, Opus 4.8, surprisingly, is still the better choice for two specific kinds of work.

  4. Guides

    Why AI sounds confident when it's wrong

    A language model's confidence reads like clean handwriting: the page stays just as neat whether the claim underneath is solid or hollow. Why right and wrong arrive in the same voice.

  5. Guides

    AI observability proves the run finished. An eval proves it was right.

    LangChain's State of Agent Engineering survey of 1,340 practitioners: 89% run observability, only 52.4% run offline evals. An offline eval is what scores whether the answer was right.

More on the site169 pages

Analyses 37

Calculators 45

Glossary 60

Independent research, built and maintained by LatentEval. We show our methods and cite our sources. See our methodology and disclaimer.