Evidence and reproducibility calculators
What has to be written down for a number to be checkable later: the card behind a result, the manifest behind a re-run, and the claim each piece of evidence is actually evidence for.
What these numbers mean
An eval number that nobody can re-run is a claim, not evidence. The instruments below are about the part that survives the run: what was recorded, whether someone else could reproduce it, and which specific claim each artifact is actually evidence for. None of them re-computes a result.
| What you have | The instrument |
|---|---|
| A result and no record of the conditions it was measured under | write the evaluation card |
| A card or a manifest, and a doubt about whether anyone could re-run it | check it against the reproducibility profile |
| A set of claims going into a deployment decision, and a pile of artifacts beside them | register each claim against its evidence |
The method these instruments apply is written out in the evidence quality framework, which defines the per-claim status vocabulary and the dimensions each claim is assessed against. This family rests on our own work on the four ways an eval number lies and how to report a judge-scored evaluation honestly.
The fork here is whether a gap is absent or recorded. A card with a null temperature and a stated reason is a different object from a card that never had the field. Only the first one tells a reader what to go and find out.
Tools in this topic
Evaluation Card Generator: Markdown and JSON on a Published Schema
Write one evaluation down so a second person can read the number correctly, or run it again. Fill the form, export a Markdown card and a JSON card that validates against a schema published here.
Instrument | Evidence and reproducibilityClaim-Evidence Matrix: What Backs Each Deployment Claim
One row per deployment claim, eleven evidence dimensions from relevance and representativeness to independence and applicability, with named rate-down reasons and residual deficits.
Instrument | Evidence and reproducibilityReproducibility Checklist and Manifest Validator
Check whether your write-up records what someone else would need to rerun your eval: a checklist against the published card schema, and a validator for a pasted or uploaded manifest.
Analyses that use these calculators
- Reliability testing
The agent reliability testing checklist your green eval skips
A pre-flight agent reliability testing checklist: six checks with pass criteria and calculators, from failure-mode coverage and statistical power to pass^k, cascade budget, and judge calibration.
- Eval statistics
How many runs a reliable eval needs to catch a regression
How many runs a reliable eval needs is a power calculation set by the regression you must catch, your target power, and the baseline pass rate. Includes a runs-needed table and the formula behind it.
- Judge reliability
Bias-correct your LLM-as-a-judge eval before reporting it
An LLM judge is an imperfect classifier, so its raw pass rate is biased. Correct it with the judge's sensitivity and specificity, then report a calibration-aware confidence interval.
Where next
- Directory | 61 calculators
Evaluation and reliability calculators
Calculators for AI agent eval statistics: confidence intervals, paired significance, repeated-run reliability, judge calibration, agreement and bias, prompt robustness, and RAG.
- Reference
Glossary
The metrics these calculators implement, defined in plain language with their assumptions.
- Glossary
Eval reproducibility
Eval reproducibility is getting the same result from an evaluation re-run on the same data and the same parameters; it breaks when uncontrolled non-determinism such as sampling temperature, an unpinned seed, or a drifting judge model moves the score while the declared inputs stay fixed.
- Analysis
The agent reliability testing checklist your green eval skips
A pre-flight agent reliability testing checklist: six checks with pass criteria and calculators, from failure-mode coverage and statistical power to pass^k, cascade budget, and judge calibration.
- Analysis
How many runs a reliable eval needs to catch a regression
How many runs a reliable eval needs is a power calculation set by the regression you must catch, your target power, and the baseline pass rate. Includes a runs-needed table and the formula behind it.
- Analysis
Bias-correct your LLM-as-a-judge eval before reporting it
An LLM judge is an imperfect classifier, so its raw pass rate is biased. Correct it with the judge's sensitivity and specificity, then report a calibration-aware confidence interval.