Who publishes this
About LatentEval
Models keep getting more capable; the systems built on them do not get more reliable at the same pace. A model can top every benchmark and still poison shared state three agents deep in a pipeline, turn one bad retrieval into a confident wrong answer, and pass every check while doing it. Capability and reliability are two different curves, and LatentEval is about the second one.
LatentEval is an independent research property on the reliability of AI agent systems. We publish meta-analyses of what the field has actually shown, original benchmarks and experiments on real-world multi-agent systems, and plain-language write-ups that make that evidence usable for the people building and relying on these systems. We show our methods, cite our sources, and report uncertainty honestly without false precision.
What we publish
Each piece answers a single question: how far a fault propagates across a multi-agent topology, whether a reported eval difference is statistically real, how much variance a benchmark actually has. The work takes three forms, and the first two are published together in Research. Meta-analyses pool what published studies have actually shown, with the inclusion criteria and effect sizes stated. Original benchmarks and experiments measure real-world multi-agent systems under controlled, repeatable conditions with statistical rigor. Plain-language write-ups, including the glossary and the guides for people building agents and for everyone else using them, make that literature readable for engineers, leads, and practitioners as well as researchers.
Our honesty contract
We hold ourselves to four standing promises, and we won't publish work that breaks one of them.
- The finding first. Every piece leads with its finding. The method, the assumptions, and the fine print sit directly below it.
- The work up front. We do not place ads or advertising trackers between you and the work.
- We show our work. Every claim names its method and cites its sources. Where we report a reliability number, we show the uncertainty behind it. For work that ages (models, harnesses, baselines), we note when we last reviewed it.
- Nothing invented. We never fabricate results, citations, or credentials. Where the evidence is thin, we say so plainly.
Independence
Independence is the point. LatentEval takes no advertising and accepts no arrangement that would give us a reason to make one system's reliability look better or worse than the numbers say. Reliability findings are only worth anything if the people publishing them have nothing riding on the result. If our funding model ever changes, we'll say so plainly on this page before it happens.
Privacy
We keep all data usage transparent and upfront. The only measurement today is aggregate analytics, described plainly on our privacy page. If what we collect ever changes, that page is updated first.
How we show our work
We believe a result you can't check isn't worth trusting. Our methodology page explains how we choose methods for each analysis, how sources are pinned to a publisher, a direct link, and a retrieval date, and how often we review published numbers as models and harnesses change. Our benchmarks and experiments are built to measure what ordinary testing misses: how failures propagate across agents, and whether a difference is statistically real. See our disclaimer for how to read a finding and the uncertainty around it.
Contact
Found a mistake or an outdated number? Email us at contact@latenteval.ai and we'll take a look.
LatentEval is independently built and maintained. Agent reliability is a fast-moving field, and we hold our own numbers to the standard we hold everyone else's: versioned methods, cited sources, and uncertainty reported honestly. When a call is close, our disclaimer explains how to read it.
About the author
I'm Srivatsa Koganti, and across more than fifteen years I've worked at the high-consequence end of applied AI and analytics, spanning all the major industries and multiple regions.
How I see the problem owes as much to two thinkers as to any ML paper I've been inspired by. From Donella Meadows, the systems lens: complex systems fail through feedback loops and the way their parts interact, not through single broken components, and the failure you can foresee is rarely the one that hurts you. From Daniel Kahneman, the cognitive lens: intelligence runs on heuristics that are efficient right up until they are quietly, systematically wrong. Modern AI inherits both at once. It reasons through learned heuristics that carry their own biases and blind spots, then runs inside multi-agent systems where those local errors couple, feed back, and cascade in ways no single-model benchmark can see. That collision, cognitive failure modes moving through systems dynamics, is where reliability actually breaks, and it is genuinely hard to predict. Mapping it is something the world will have to master to make AI dependable and safe. I don't expect to solve that alone; I intend to contribute a real, measured piece of it, because the value AI can create is capped entirely by how reliably and predictably we can trust it.
I've led complex AI and ML systems into and through production at Fractal and Axtria, two of the firms that actually ship enterprise AI at scale, delivering for some of the world's largest technology, consumer, and life-sciences companies, and building multiple data platforms from idea to multi-million ARR. The statistical discipline comes from a different lineage: postgraduate work in Computational Data Science at the Indian Institute of Science, and years delivering data-science and ML products for the C-suite and senior leadership of industry leaders inside McKinsey and BCG, with earlier enterprise work at Franklin Templeton and UnitedHealth Group. That pairing is what this work runs on: reliability as distributed systems taught it, where what matters is how far a failure propagates and whether it's contained, measured honestly rather than asserted.
A word on independence. LatentEval sells nothing and takes no sponsorship: no vendor I'm quietly steering you toward, no reason to make any system look better than it is. Much of what passes for "reliability research" is observability marketing with footnotes. This is not that. Finding first, strong baselines, every claim carrying a number and the statistical honesty behind it, and when a widely believed result fails to replicate, that gets published too. The work is built to be checked, which is what makes it worth citing. If you build agentic systems, evaluate them, or have to vouch for them to someone who will hold you responsible when they fail, this is written for you.