Evaluation design calculators
Everything that happens before a run: what the system is for, which failures matter, what a case looks like, who grades it, and what the result will and will not prove.
What these numbers mean
Most of what makes an eval wrong is decided before anything runs. Which population the cases come from, whether the expected answer is a reference or a judgment, which failures get their own count, and which question the number will be asked to settle are all design choices. A confidence interval computed afterward cannot recover any of them. The instruments below are for that stage. None of them runs a suite and none of them calls a model.
| What you are trying to work out | The instrument |
|---|---|
| What this system is actually for, and what that changes about testing it | profile the context of use |
| What the whole evaluation will test, written down before results exist | write the evaluation blueprint |
| Which test cases a named failure concern actually calls for | map concerns to test cases |
| Which parts of the testing surface have no cases at all | build the coverage matrix |
| Which slices to plan for, and how many cases each one needs | plan the slices before the set exists |
| Whether to score both systems on the same cases or different ones | choose paired or independent |
| How many times to run each case before the spread is readable | plan the repeated runs |
| What the design will cost before you commit to it | estimate the eval budget |
| Whether to score outputs one at a time or against each other | choose the scoring protocol |
| Who or what should grade this, and which metric answers the question | choose the evaluator and the metric |
| What the evidence you collected actually licenses you to claim | check what the evaluation proves |
| What a case looks like for the task family you are testing | start from a task-family blueprint |
| How many distinct items the set needs, per slice and in total | size the golden set |
| What "good enough to ship" means, in writing, before the run | set the acceptance criteria |
This family rests on our own work on the four ways an eval number lies, which eval number needs which test, and what a single pass rate hides about run-to-run consistency.
The fork people get stuck on is whether to write more cases or to write different ones. Coverage of the failures you already named is a different property from coverage of the population you actually serve, and only the first is visible in a grid. The coverage matrix shows that. The context profile tells you whether the second one is the property you should be worried about.
The second fork is the expected-answer column. A set that mixes a reference answer, a rubric and an expected end state in one column and then applies one grader to all three produces a number whose meaning changes row by row. Declaring the oracle type per case is what the case schema is for, and the evaluator navigator lists which graders the declared oracle admits, and the schema validator checks each row against that list.
Tools in this topic
Context-of-Use Profiler: Turn a System Description into Evaluation Consequences
Describe what your AI system does, who it affects and how it fails, and get the failure families, case classes, slices and evidence expectations those answers demand.
Instrument | Evaluation designEvaluation Blueprint Builder: Write the Argument Before You Run the Eval
State each claim your evaluation will make, name the evidence that would support or refute it, and write the inference step connecting the two, all before any results exist.
Instrument | Evaluation designRisk-to-Test Mapper: Which Failures Need Which Cases
Name the failures you care about, then see the test-case shapes, slices, evaluators and metrics each one needs. Every mapping says what it cannot establish. The mapping is a LatentEval synthesis.
Instrument | Evaluation designEvaluator and Metric Navigator: Which Grader, Which Number
Answer four questions about oracle type, grader, system count, and failure mode. Get the admissible evaluator, inadmissible ones and why, the metric family, and a route to the tool that runs it.
Instrument | Evaluation designWhat Does This Evaluation Prove?
Select the evidence your evaluation produced, see which claims it supports, which it cannot, and the inference error behind each gap. Each unsupported claim names the instrument that would close it.
Instrument | Evaluation designJudge Validation Report Builder: Report, Version Card, Drift Plan
Fill one form and keep three documents: a validation report showing your LLM judge was checked against people, a version card pinning the prompt and settings, and a drift plan with a date on it.
Instrument | Evaluation designRelease Gate Designer: Go/No-Go Rules and Decision Log
Write the rule that turns your eval figures into PROMOTE, HOLD or ROLLBACK, get the verdict plus the rule that produced it, and keep a dated decision log and memo in your browser.
Instrument | Evaluation designEvaluation Coverage Matrix: Which Cells Have No Cases
Put what you test on one axis and what you test it across on the other. The grid shows which cells have cases, which are empty, which are too small to carry an interval, and what each gap costs you.
Instrument | Evaluation designAcceptance Criteria Builder: Set the Bar Before the Run
State each pass/fail criterion, the risk axes it covers, the reasoning behind the threshold and who set it. Get the sample size each one needs and export the rules for the release gate designer.
Instrument | Evaluation designGolden Set Blueprint Builder: Seed a Test Set From a Task Family
Choose a task family, load its case taxonomy and starter rows, then edit, add and export. Every row carries an oracle type and a provenance tag, and the export validates against the eval-set schema.
Instrument | Evaluation designClaim-Evidence Matrix: What Backs Each Deployment Claim
One row per deployment claim, eleven evidence dimensions from relevance and representativeness to independence and applicability, with named rate-down reasons and residual deficits.
Analyses that use these calculators
- Judge reliability
Kappa thresholds for LLM judges, and who published each one
Five published kappa bands from four sources, side by side, each with the author who wrote it, the date we read it, and the coefficient it was written for. They are conventions, and they disagree.
- Reliability testing
The agent reliability testing checklist your green eval skips
A pre-flight agent reliability testing checklist: six checks with pass criteria and calculators, from failure-mode coverage and statistical power to pass^k, cascade budget, and judge calibration.
- Judge reliability
The LLM-judge bias checklist that gates your ranking
The LLM-judge bias checklist is eight pass/fail gates you run before trusting a ranking. Each gate pairs a detection test with a numeric pass line and the calculator that computes it.
- Reliability testing
Agentic AI testing: the five dimensions one eval run cannot reach
How to design the test rather than pick a metric: run count and power, run-to-run variance, fault injection across topologies, and the protocol that turns five dimensions into a release decision.
- Reliability testing
How to measure agent reliability past a single pass rate
How to measure agent reliability with metrics that capture the consistency a single pass rate cannot: pass@k versus pass^k, a reliability@k suite aggregate, and a confidence interval on every rate.
Where next
- Directory | 61 calculators
Evaluation and reliability calculators
Calculators for AI agent eval statistics: confidence intervals, paired significance, repeated-run reliability, judge calibration, agreement and bias, prompt robustness, and RAG.
- Reference
Glossary
The metrics these calculators implement, defined in plain language with their assumptions.
- Glossary
Judge calibration (LLM evals)
Judge calibration is the correspondence between the confidence an LLM judge attaches to a verdict and how often verdicts carrying that confidence turn out correct, measured against held-out human labels rather than assumed from the judge's own scores.
- Glossary
Rubric drift (LLM judges)
Rubric drift is the movement of an LLM judge's effective scoring standard while the rubric text it is sent stays fixed, so two scores produced under the same rubric no longer sit on the same scale. The instrument changed between the measurements.
- Analysis
Kappa thresholds for LLM judges, and who published each one
Five published kappa bands from four sources, side by side, each with the author who wrote it, the date we read it, and the coefficient it was written for. They are conventions, and they disagree.
- Analysis
The agent reliability testing checklist your green eval skips
A pre-flight agent reliability testing checklist: six checks with pass criteria and calculators, from failure-mode coverage and statistical power to pass^k, cascade budget, and judge calibration.
- Analysis
The LLM-judge bias checklist that gates your ranking
The LLM-judge bias checklist is eight pass/fail gates you run before trusting a ranking. Each gate pairs a detection test with a numeric pass line and the calculator that computes it.
- Analysis
Agentic AI testing: the five dimensions one eval run cannot reach
How to design the test rather than pick a metric: run count and power, run-to-run variance, fault injection across topologies, and the protocol that turns five dimensions into a release decision.
- Analysis
How to measure agent reliability past a single pass rate
How to measure agent reliability with metrics that capture the consistency a single pass rate cannot: pass@k versus pass^k, a reliability@k suite aggregate, and a confidence interval on every rate.
- Analysis
Bias-correct your LLM-as-a-judge eval before reporting it
An LLM judge is an imperfect classifier, so its raw pass rate is biased. Correct it with the judge's sensitivity and specificity, then report a calibration-aware confidence interval.