Glossary
Effect size (eval deltas)
Effect size is the magnitude of a difference between two eval results, measured on a scale that holds still when the run count changes: on a pass/fail suite, the gap between two pass rates in percentage points, reported with an interval on the delta itself.
Effect size is the magnitude of the difference between two eval results, measured on a scale that holds still when you change the number of runs. On a pass/fail suite that is normally the raw difference between the two pass rates in percentage points, computed over the same task set: a candidate at 87% against a baseline at 80% is an effect size of 7pp. It answers how big the gap is and says nothing about how much confidence the gap has earned, which is the separate reading a test on the delta supplies.
A difference between two proportions inherits the uncertainty of both, so the delta needs an interval of its own. Checking whether the two arms’ confidence intervals overlap is a weaker test than putting an interval directly on the difference, and on a paired design, where both systems scored the same items, it also discards the pairing that made the comparison tight.
Two of our own pre-registered studies fixed an effect size before any data arrived, and this page rests on both. The routing study on short machine-checkable tasks declared a separation bar of at least 15 percentage points together with disjoint 95% Wilson intervals, then measured a Fable-minus-Opus gap of 17.9 points against Fable at low effort and 3.6 points in its favor at extended-high. The low-effort magnitude cleared 15 points on a Fable sample of seven non-refused trials, the two intervals overlapped anyway, and the study published an inconclusive null. The within-Anthropic three-way set a 2.0-point winning threshold in advance, came back 1.13 points apart at 93.8 against 92.7, and reported a tie with no order inside it. Name the delta you would act on before the run, and the number that comes back has a decision waiting for it.
How to calculate effect size on an eval delta
Score both arms on the same task set, take each arm’s pass rate over its own denominator, and subtract the baseline from the candidate. Report three things together: the point delta, an interval on that delta, and both denominators, because a 7-point gap over 30 items and a 7-point gap over 3,000 items are the same effect size carrying wildly different evidence. Denominators are where evals leak, since refusals, timeouts and truncations quietly shrink the set a rate is computed on, which is why answer coverage, how much of the intended sample survived to be scored, belongs beside any delta you publish.
Where the two runs scored identical items, recover the interval from the paired table with the McNemar calculator. Where they did not, put a Wilson interval on each rate and accept that an unpaired comparison is the blunter instrument. A statistic with no closed form, such as a delta between two weighted composites, gets its interval from resampling the runs instead. Standardized forms exist, Cohen’s h for proportions and Cohen’s d for continuous scores, and they earn their place when you pool results measured on unlike scales, which a suite running one fixed rubric rarely needs.
Instrument it where the two arms meet: a per-item results table keyed on a stable item id, so the paired counts can be recovered later. Record the condition alongside the number, since a delta measured at one effort level, one temperature and one prompt phrasing is a claim about that cell and no other, and a re-run under a fresh seed can move it.
Effect size vs statistical significance
Statistical significance is a statement about compatibility with chance: assuming no true difference between the arms, how improbable a difference at least this large would be. It is computed from the effect size and the sample size jointly, so one delta drifts in and out of significance as the arms collect more runs.
Run count moves significance, never the delta itself.
Either one can move while the other holds still. A large effect misses significance when the denominator is small, which is what happened to that 17.9-point low-effort gap: Fable’s own capability rate sat at 71.4% on seven non-refused trials, 95% CI 35.9–91.8%, and an interval that wide cannot exclude zero however far apart the two points land. A negligible effect clears significance when the denominator is large: on a suite of tens of thousands of items, a fraction of a point becomes hard to attribute to chance while sitting far below anything you would roll a release back for. Report both, and say which one your decision turns on. Significance speaks to whether the difference is there at all, and effect size to whether it is worth acting on.
Effect size vs statistical power
Statistical power is the probability that an eval detects a real difference of a stated size, given the run count, the baseline rate and the significance threshold in use. The effect size is one of its inputs, which is the whole of the relationship: you name the regression you refuse to miss, and the power arithmetic returns how many runs buy a stated chance of catching it. The sample size and power calculator runs that step in either direction.
They separate on which one you choose and which one you are handed. Effect size is a property of the two systems being compared, and no amount of instrumentation shifts it. Power is a property of the instrument, and you raise it by adding runs, pairing the design, or cutting seed-to-seed spread. Our sizing page carries the price list for that trade: from a 90% baseline at 80% power, catching a 10-point drop takes roughly 70 runs while a 2-point drop takes about 1,500, since halving the effect roughly quadruples the count. Set the effect first and read the runs off it. Picking a run count you can afford and finding out afterwards which regressions it was blind to is the same calculation run backwards, and it tends to end with an eval that could only ever have caught a collapse.
Where the psychology conventions stop applying
Effect size is borrowed vocabulary, and the field that owns it is behavioral science, where a standardized statistic and a table of labels both do real work. Cohen’s benchmarks are what most people carry across: small, medium and large pinned to d values of 0.2, 0.5 and 0.8, which Daniël Lakens’s open statistics textbook calls arbitrary and advises against using at all. That criticism lands on the labels rather than on the standardization, which earns its keep in a discipline pooling incommensurable instruments across hundreds of studies.
An eval delta sits in a different position from any of that. The outcome is usually binary on a task set you control, both arms ran the same items, and the number feeds a ship-or-roll-back decision with a cost attached on each side. Percentage points are already the natural unit, so standardizing buys you little, and the threshold worth having is the smallest regression your team would actually act on. What earns the decision instead is a bar declared in the units of the metric before the run and reported against afterwards, which is what both of our studies did: 15 percentage points with disjoint intervals in the routing work, 2.0 index points in the three-way. A conventional label imported from another field would have settled neither one.
Effect size and statistical power are computed from each other, so a report that publishes one and drops the other leaves its reader unable to separate a genuine equivalence from a thin sample. Which statistics an eval owes its reader, and which calculator answers each of them, is collected in the eval statistics hub. Write the delta, its interval and both denominators into the same line of the report, and a reader can redo the rest of the arithmetic without asking you for it.