LatentEval
Eval statistics

Eval statistics: which number needs which test

The statistics an agent eval rests on, routed by the question in front of you: sizing before the run, the interval on a rate, a paired test on a delta, and agreement on the labels.

Part of AI agent evaluation that follows the whole trajectory

Eval statistics

In brief

5 POINTS
  • Five questions sit behind an eval number, and each one needs a different instrument to answer.
  • Sizing happens before the run; an eval too small to see a regression cannot report one.
  • A pass rate quoted without an interval hides how much of a gap between two systems is sampling noise.
  • Our three-way benchmark published a tie because the measured gap fell under a threshold fixed in advance.
  • Agreement between a judge and a human reference is a separate measurement from the pass rate itself.

Our three-way Claude reliability benchmark scored Claude Fable 5 at 93.8 and Claude Opus 5 at 92.7 on a composite index, then declined to say which one won. The rule was fixed before the first call went out: a gap under 2.0 points on that index is a tie. The gap came in at 1.13 points on the unrounded index.

Nothing in the data stopped us from writing that Fable 5 edged ahead. The direction was there, the decimals were there, and a reader would have taken it on trust. A threshold set in advance is why that sentence never reached the page.

The same discipline shows up from the other direction in our pre-registered routing and refusal-tax study, where Fable 5’s capability rate at low effort rests on seven non-refused calls: 71.4%, 95% CI 35.9–91.8%. As a point estimate it reads like a comfortable pass. The interval runs from worse than a coin flip to near certainty, and which task types justify reaching for the pricier tier is the decision that width has to survive.

Settle the threshold, the run count and the reporting rule before the eval runs, then publish every number beside the interval or the test that licenses it. An eval that skips that order still returns a number, and the number will look exactly like one you could act on.

Those two runs are worked examples of five separate statistical questions. The five collapse easily into a single demand for rigor, which then gets answered with whichever test the harness already has wired in. This page pulls them apart and sends each one to the page and the calculator that handles it.

An eval number estimates something you never measured

A pass rate is computed on the suite someone happened to write, scored over the attempts the harness happened to sample. Evan Miller’s statistical treatment of evals begins there, and it treats the questions in a suite as a sample drawn from an unseen super-population. It opens by observing that the eval literature “has largely ignored the literature from other sciences on experiment analysis and planning” (arXiv preprint, November 2024). Two suites of the same size, drawn from the same pool of tasks, land on different scores. That spread is a property of the measurement, which is why a bare percentage is an incomplete report of what an eval found.

Before any of this bites, work out which of the four eval instruments produced your number. An offline benchmark, an LLM judge, a human review panel and a production metric each answer a different question and each carry a different kind of uncertainty, so the same figure means four different things depending on where it came from.

Five questions follow, in roughly the order a working eval meets them. Four of the five have a calculator on this site, and the label question takes two of them. The table carries a sixth row for repeat-attempt consistency, which is the one piece of the fifth question that a calculator can answer. Every calculator here has a boundary it won’t cross, so the last column says what each one leaves to you.

TABLEShow full table (6 rows)Showing full table (6 rows)
The question in front of youThe calculatorWhat it will not answer
How many runs before this comparison could see a regression at all?Sample size and powerWhether the regression size you sized for is the one worth rolling back over
How wide is this pass rate, really?Wilson and Clopper-Pearson intervalsWhether the tasks in the suite were the right tasks
Did this delta survive on the same items?Paired McNemar testHow large the difference is in points, which is a separate reading
How often does the same task pass on repeat attempts?reliability@k and pass^kWhy the failures cluster on the tasks they cluster on
Does the judge agree with a human on the same labels?Cohen’s kappa with an intervalWhether judge and human are wrong in the same direction
How far does the judge’s own error bend the rate you publish?Judge bias correctionWhich bias produced the error

The table routes the arithmetic. The rest of this page routes the reasoning, because every one of those calculators will happily return a number for inputs that had no business being combined.

The run count is settled before the eval runs

Three inputs fix how many runs a comparison needs: the size of the regression you would actually roll back for, the probability you want of catching it, and your baseline pass rate. The first two are judgment calls about what a regression is worth to you, and the third comes off the baseline eval you already run. A fourth input sits behind the arithmetic and rarely gets said out loud: the significance level, which the sizing formula holds at the conventional 0.05 until somebody moves it. How many runs a reliable eval needs does that arithmetic and carries the table of runs required at each drop size, so it is the page to open when the question is sizing.

The probability half of it has an entry of its own. Statistical power for an eval design is the chance the suite catches a regression that is really there; the entry covers how power is computed and why an underpowered suite returns nulls that carry no information. It also tracks the conventional 80% target back to where it was set. Jacob Cohen proposed 0.80 in 1992 as a convention for general use in the behavioral sciences (Cohen, 1992, peer-reviewed), balancing the risk of missing a real effect against the sample a higher target would demand. Taken with the conventional 5% false-positive rate, that pair treats a false positive as four times as serious as a missed effect. A release gate keeps a different ledger, and a team that inherits the 4:1 ratio has let another field’s cost structure set its rollback risk.

An eval sized after the fact cannot be repaired by analysis. Once the runs are spent, the only remaining moves are to spend more of them or to widen the difference you are willing to call.

How wide is the number you just published?

A pass rate is a proportion, and a proportion carries an interval whose width falls out of the run count. Why an eval confidence interval belongs beside every rate sets out the closed-form case, where k passes out of n attempts have a Wilson or exact interval waiting for them. The seven-call example at the top of this page is what a closed-form interval looks like at small n, printed at its full width.

Ladder of Wilson 95% confidence intervals on a 0 to 100% pass-rate axis: the measured row, seven calls at 71.4%, spans 35.9 to 91.8 and crosses the 50% coin-flip line; computed rows at 15, 30, 60, 120 and 240 calls hold the same rate and narrow to 65.4 to 76.8.
A 71.4% pass rate over seven non-refused calls reads like a comfortable pass as a point estimate; its Wilson 95% interval runs from 35.9 to 91.8, worse than a coin flip to near certainty. The rows below hold the rate fixed and recompute the interval at larger run counts: the width is a property of the run count, settled before the eval runs. Top row: measured values from our model routing and refusal-tax study. Lower rows: Wilson 95% intervals computed from the formula at the same 71.4% rate; derivations, not observed eval results.

Plenty of eval statistics have no closed form at all.

A weighted composite across dimensions, a mean of per-task rates, a suite-level consistency score: none of them come with a textbook standard error. Bootstrap resampling the scored runs to get an interval anyway is the general-purpose answer, and it is what a harness falls back on when the statistic it is reporting has no formula sitting behind it.

Interval width has one trap in it. Re-running the same twelve tasks four hundred times narrows the interval beautifully and tells you nothing new about the next twelve tasks. Width responds to attempts. Whether the suite represents the work is answered somewhere else entirely, and so is the question of what a gap between two of these rates is worth.

A delta is a hypothesis until a paired test scores it

A delta arrives as two numbers and a subtraction, and the subtraction is the part nobody puts an error bar on. Is your eval difference statistically significant? works the paired case properly. It scores the items where the two systems disagreed instead of the two headline rates, which is what makes McNemar the right instrument when both systems ran on the same tasks.

Significance and magnitude are two different readings, and a report carrying one without the other leaves its reader unable to act. How large a pass-rate difference actually is covers the magnitude side: a 7-point gap in pass rate is a 7-point gap whether it came from 40 runs or 4,000. The interval belongs on the delta itself, where the subtraction happened. A statistically significant 0.4-point improvement in pass rate is a real finding and still not a reason to ship.

Our own benchmark could not put an interval on its delta. The composite index publishes no interval envelope, so there was no error bar to place on the gap and no paired test behind the call. A 2.0-point winning threshold was fixed before the first call went out instead, the observed gap landed under it, and the study published a tie with no order inside it. A bar set in advance answers a narrower question than a paired test does, and it removes the same temptation: to discover afterwards that the bar was wherever the result landed.

Number line of the composite index from 90 to 96: Opus 5 at 92.7 and Fable 5 at 93.8 with their 1.13-point gap marked, inside a gold tie band running from 92.7 to the 94.7 winning threshold that was fixed before the run; the study published a tie.
The measured gap between the two composite indices is 1.13 points against a winning threshold of 2.0 fixed before the first call went out. Fable 5's score lands inside the tie band, so the study published a tie and no order inside it: the verdict came from the pre-registered rule alone. Teal values are measured results from our Fable 5 vs Opus 5 vs Opus 4.8 reliability benchmark; the gold band is the pre-registered decision rule, not a measurement. Index window 90 to 96 shown.

The labels came out of an instrument too

Every number so far assumes the pass and fail labels are correct. When a model assigned them, that assumption is a second measurement with its own error rate, and it needs characterizing before its output enters any of the arithmetic above.

Cohen’s kappa for judge-against-human agreement is the usual starting point for two raters on pass/fail labels, chance-corrected so that a lopsided suite cannot manufacture agreement out of base rates. It also has a documented failure mode, the first of the two kappa paradoxes Feinstein and Cicchetti named in 1990 (Feinstein and Cicchetti, 1990, peer-reviewed), where raters match on nearly every item and kappa still collapses because the suite is lopsided. The entry works that case through. When the design outgrows two raters or a binary label, Krippendorff’s alpha across any number of raters and any scale handles graded rubrics and missing labels, and the two entries split the space between them.

Agreement is one axis of three. Is your LLM-as-a-judge reliable? separates agreement from calibration and from bias, each with its own test. Once a judge has been characterized on all three, bias-correcting a judge eval before you report it is the workflow that turns it into a publishable rate carrying both sources of uncertainty.

The bias axis has the most machinery behind it. The taxonomy of judge biases and the tests that catch each is the map, and the gate order you run before publishing a ranking is the procedure that walks it. Start at the swap-consistency test for a pairwise judge, which asks only that you re-run each comparison with the two candidates in the other order and count how often the verdict flips. The term-by-term index of the vocabulary sits in the judge bias vocabulary.

Which part of the spread could you actually shrink?

An interval tells you how wide the number is. It doesn’t say which knob would narrow it, and those are two problems with different price tags. Variance decomposition splits eval spread into the sources that produced it and names the facets: which tasks the suite happened to contain, how the model sampled tokens on each attempt, which judge scored the output, and what the harness held fixed. More attempts per task shrink one facet cheaply. Writing more tasks or replacing the judge is a different budget. That cheap facet is also the only one here with a calculator behind it: how often the same task passes on repeat attempts is the attempt-to-attempt spread measured directly, and the other three facets are read off the design rather than computed.

Two neighboring properties sit alongside it and are routinely confused with it. Whether a re-run lands where the last one did is a property of repetition, and a suite can be perfectly reproducible while most of its variance hides in the tasks nobody sampled. Scores computed over only part of the instrument is the accounting for the other case, where a chunk of the eval never returned a result at all.

That last one is why our benchmark’s tie carries a second caveat. Opus 5’s index was computed over 77.3% of the instrument weight, because the provider blocked every request on the one task deep enough to separate the two models. Fill that missing weight with zeros and its index is 71.6; fill it with full marks and its index is 94.3. That range is a bound on where the true index could sit, and more calls on the tasks that did run leave it exactly as wide. A blocked call is a missing observation, and the share of items that came back with an answer at all belongs next to any score computed on the survivors.

Where the statistics stop and validity begins

Everything on this page assumes the suite measures the thing you care about. No interval, test or agreement coefficient can rescue a benchmark that scores the wrong construct, and whether the suite measures what its name claims is the prior question all of them depend on. A perfectly powered eval with tight intervals and an audited judge, run on the wrong task, is a precise measurement of nothing in particular.

For agent systems the construct problem gets sharper, because a final-answer score cannot separate a correctly reasoned path from one that stumbled early and recovered by luck. AI agent evaluation across the whole trajectory rather than the endpoint is the pillar for that, and it is where the statistics here get applied to multi-step runs.

The wider orientation, including four measurable ways an eval number can mislead you and the fix for each, lives in which eval methods to trust and where they lie. Start there if the question is what an eval certifies at all rather than which test to run.

The whole cluster reduces to a checklist a release template can hold. Decide the regression size and the threshold before the run, size the run to that, publish every rate with its interval, publish every delta with a paired test and its magnitude, publish the coverage the score was computed over, and characterize any model that assigned a label before you quote what it scored. Six lines. Every one has a page on this site behind it, and four of them have a calculator.

Sources

  1. Claude Fable 5 vs Opus 5 vs Opus 4.8 reliability benchmark Retrieved
  2. Model routing and the refusal tax: a pre-registered study Retrieved
  3. Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations Published
  4. A Power Primer, Psychological Bulletin 112(1), 155-159 Published
  5. High agreement but low Kappa: I. The problems of two paradoxes, Journal of Clinical Epidemiology 43(6), 543-549 Published