LatentEval
Topic

LLM eval sample size and planning calculators

5 calculators | 5 analyses

Everything here is decided before the run: whether both models see the same items, how many items the set needs, how many repeats each item gets, and what that adds up to in tokens and dollars.

The method

What these numbers mean

Every number on this page is settled before a single case is scored. A plan that skips them still produces a result, and that result carries whatever precision the plan happened to buy, which is usually less than whoever reads it assumes. The four calculators here turn the four planning questions into counts you can defend: which design the comparison uses, how many cases the set holds, how many times each case runs, and what the whole thing costs to execute.

Work across the table in that order. The design choice comes first because it changes what every later count means. The case count and the repeat count multiply together, so neither one is readable on its own. Cost sits last because it takes both of the others as inputs, and it is the row that most often sends people back to revise the first three. One more card sits in the grid below, cross-listed from pass-rate-statistics: the live sample size and power calculator, which sizes the same run count from a classical two-arm or paired-design formula for a reader who wants that approach instead of this family’s variance split.

What you are decidingThe instrument
Whether the two systems being compared see the same casessettle the design before anything else
How many cases the set has to holdsize the set from the drop worth catching
How many times each case should be runsplit the variance, then set the repeats
Cases in the pilot that were not all run the same number of timesweight the ragged pilot instead of trimming it
What the plan will cost in tokens and in review hoursprice it before you commit to it
A run that already happened and a number that needs defendingtake it to the after-the-run instruments

The repeat guidance here rests on our own measurement of how far a suite’s score moves between runs that changed nothing: how many times a suite has to run before its number holds still. The wider map of eval statistics picks the thread up after the run finishes, which is exactly where this page stops.

The choice people get stuck on is whether to spend on more cases or on more repeats. More cases buy down case-to-case variance, more repeats buy down run-to-run variance, and the variance split is what tells you which of the two you are actually paying for. Budget is what turns that into a real decision rather than a preference, so price both plans before committing to either one.