Topic
Eval statistics
The statistics under an eval result: sample size, confidence intervals, the tests that separate a real difference from noise, and what an eval can certify.
Research
- Analysis
Benchmark contamination: what you can check and what you cannot
Benchmark contamination is two problems under one name. Leakage inside your own splits is checkable in a browser in minutes. Corpus-level contamination is a research problem.
- Analysis
Eval statistics: which number needs which test
The statistics an agent eval rests on, routed by the question in front of you: sizing before the run, the interval on a rate, a paired test on a delta, and agreement on the labels.
- Study
Claude Fable 5 vs Opus 5 vs Opus 4.8 reliability benchmark
The full three-way benchmark behind our builder guide. Claude Fable 5, Claude Opus 5 and Claude Opus 4.8 on identical tasks, seven areas scored, every count and caveat published.
- Study
Model routing and the refusal tax: a pre-registered study
On short, checkable tasks we measured no Opus 4.8-to-Fable 5 capability separation at either effort we ran, and the premium tier’s refusal ‘rescue’ silently served the cheaper model on 20 of 28 calls.
- Analysis
Is your eval difference statistically significant?
Two eval runs a few points apart. Separate a real gain from run-to-run noise with a paired McNemar test on the same items, then read what a significant result does and does not license you to claim.
- Analysis
How many runs a reliable eval needs to catch a regression
How many runs a reliable eval needs is a power calculation set by the regression you must catch, your target power, and the baseline pass rate. Includes a runs-needed table and the formula behind it.
- Analysis
AI agent evaluation that follows the whole trajectory
AI agent evaluation breaks when it scores the final answer and skips the path. Evaluate the trajectory, catch early-step corruption, and report pass rates with intervals.
- Analysis
What LLM evals are, and what each type can certify
LLM eval covers four instruments: offline benchmark, LLM-as-judge, human, and online, each answering a different question, plus the benchmark-vs-product line and the rigor behind a trustworthy score.
- Analysis
LLM evals: which methods to trust and where they lie
LLM evals report whether a model passed. Whether that score is valid is a separate question. This hub maps the eval methods and the four ways an eval number lies, each routed to its fix.
Guides
- Guide
Evaluation case schema, field by field
The case profile published atop the eval-set schema: what each field does, why the oracle type gates the evaluator, what is left out on purpose, and how rows round-trip.
- Guide
Evidence quality for AI claims: eleven dimensions, no score
Eleven dimensions for appraising evaluation evidence, named rate-down reasons for each, and a five-part argument against collapsing them into a composite number.
- Guide
Golden set blueprint: RAG and knowledge assistants
Seven case families for a RAG assistant, five slice families, oracle types and evaluator guidance per family, downloadable starter rows, and a pharma overlay.
- Guide
Golden set blueprint: structured extraction
Six case families for document extraction, counting omission and hallucination separately and reporting both per-field and per-record metrics.
- Guide
Golden set blueprint: summarization
Six case families for summarization: source contradictions, unsupported additions, and salient omissions counted separately, with evaluator granularity tested as a variable.
- Guide
OpenAI Evals is winding down. The alternatives skip the statistics.
OpenAI is deprecating its hosted Evals platform and steering users to Promptfoo, which it now owns. What OpenAI Evals, DeepEval, Ragas, TruLens, and Promptfoo each do, and what porting costs you.
- Guide
The prompt wording is a hyperparameter you never swept.
Rewording the same task swings a model's pass rate: format, option order, even a 'please'. A one-phrasing eval samples one point from a spread you never measured. Pin the prompt and measure it.
Terms
-
Answer coverage
Answer coverage is answers returned over requests sent. Publishing it beside any rate computed on those answers lets a reader see how much of the intended sample the rate actually rests on.
-
Bootstrap resampling (eval intervals)
Bootstrap resampling estimates the uncertainty of an eval statistic by resampling the scored runs with replacement, recomputing the statistic on each draw, and reading the spread of those values as its sampling distribution. It supplies an interval where no closed-form standard error exists.
-
Capability tier (model routing)
Capability tier is the band a router sorts a model into, ordered by how much task competence its vendor claims it delivers. The ordering is published as a product hierarchy, so whether a given boundary changes your results is a question only a paired eval on your own tasks can settle.
-
Construct validity (benchmarks)
Construct validity is the degree to which a benchmark measures the specific capability it claims rather than a proxy a system can score high on without having it; a benchmark is construct-valid only when its top score cannot be earned without the capability it advertises.
-
Coverage conditioning
Coverage conditioning is the dependence of a published rate on which requests came back with an answer, and it bites when membership of that answered subset correlates with the property the rate is meant to measure.
-
Effect size (eval deltas)
Effect size is the magnitude of a difference between two eval results, measured on a scale that holds still when the run count changes: on a pass/fail suite, the gap between two pass rates in percentage points, reported with an interval on the delta itself.
-
Eval confidence interval
An eval confidence interval is the range a procedure produces that, across repeated runs of a suite, brackets a metric's true value a stated fraction of the time (say 95%); its width combines a task-set term (closed-form binomial, or bootstrap) with the seed-to-seed spread, which one run omits.
-
Eval reproducibility
Eval reproducibility is getting the same result from an evaluation re-run on the same data and the same parameters; it breaks when uncontrolled non-determinism such as sampling temperature, an unpinned seed, or a drifting judge model moves the score while the declared inputs stay fixed.
-
Fallback chain (model routing)
A fallback chain is the ordered list of models a router tries for a single request, moving to the next entry each time the one before it declines or fails, and ending at the first model that returns an answer or at the end of the list.
-
Model router
A model router is the component that decides which model handles each incoming request, choosing once per request and before the call is dispatched, so one application can spread its traffic across an expensive tier and a cheap one.
-
Refusal rate (LLM models)
Refusal rate is the share of requests a model declines to answer on policy grounds, measured over requests sent rather than answers returned. A refusal arrives as a normal response with stop_reason set to refusal, so it never touches an error rate.
-
Statistical power (eval design)
Statistical power is the probability that an eval reports a significant difference when a regression of a stated size is genuinely present, settled before the run by the drop you would act on, the item count, the score variance, and the false-positive rate.
-
Variance decomposition (eval runs)
Variance decomposition splits the spread in an eval score into the sources that produced it: which tasks the suite happened to contain, how the model sampled tokens on each attempt, which judge scored the output, and what the harness held fixed between runs.
Calculators
- Calculator
Paired vs Independent: Choosing an Eval Comparison Design
Six questions give you the design, the test that fits it, the rules that fired, and what the branch you did not pick costs in cases and model runs.
- Calculator
pass^k, pass@k and reliability@k Estimator: Chance That All k Runs Pass
Estimate per-task pass^k and suite-level reliability@k from repeated eval runs, with Wilson and Student-t confidence intervals and the unbiased pass@k capability counterpart.
- Calculator
Golden Set Size Planner: How Many Examples an Eval Set Needs
Size a golden set from a baseline pass rate, a margin, and the drop worth catching: items per slice, a total, and a verdict on the set you already have.
- Calculator
Wilson and Clopper-Pearson Pass-Rate Confidence Interval Calculator
Turn an eval pass rate (k of n) into a defensible confidence interval: Wilson score, Clopper-Pearson exact, and Wald normal bounds side by side.
- Calculator
Eval Sample Size and Power Calculator
Find how many eval runs you need to detect a pass-rate drop at a target power, across two-arm, fixed-baseline, and paired McNemar designs.
- Calculator
Repeated-Run Variance Planner: How Many Trials per Eval Case
Split your eval's spread into run-to-run and case-to-case variance, then read the trials per case, the case count that implies, and the protocol to freeze.
- Calculator
Eval Budget Calculator: What an LLM Evaluation Run Will Cost
Price an eval plan before you run it: cases by variants by trials by judge calls, plus the human review hours it needs, from your own rates.
- Calculator
Judge Validation Report Builder: Report, Version Card, Drift Plan
Fill one form and keep three documents: a validation report showing your LLM judge was checked against people, a version card pinning the prompt and settings, and a drift plan with a date on it.
- Calculator
McNemar Test Calculator for Paired Eval Runs
Run McNemar's paired significance test on two eval runs scored on the same items, with exact and chi-square p-values and a CI on the pass-rate difference.
- Calculator
Harness Migration Score Continuity Checker
Separate what the harness moved from what your system moved. Feed the overlap bridge and your eval results to see the harness offset, the system-attributable change and a restated baseline.
- Calculator
LLM Eval A/B Comparator: Paired Difference with Bootstrap CIs
Run an LLM eval A/B test on paired graded scores: mean score difference with a bootstrap confidence interval, paired effect size, and win/tie/loss evidence.
- Calculator
Judge Agreement Tracker: Log Kappa Over Time
Keep a dated log of one judge's agreement with human labels: the value, the prompt version and set behind it, the published band it falls in, the change since the last run, and the next check date.
- Calculator
Benchmark Rank Uncertainty Calculator: Score Intervals and Rank Ranges
Put error bars on an LLM benchmark leaderboard. Paste scores with item counts, correct-of-total, or standard errors and see which adjacent ranks are a statistical tie.
- Calculator
Eval Dataset Schema Validator: Golden-Set Checks and Export Profiles
Check an eval golden set against a neutral schema in your browser: findings by row and line, duplicate ids, field completeness, plus a template and exports for four eval tools.
- Calculator
Exact Duplicate Row Checker for Eval and Golden Datasets
Find byte-identical rows in an eval or golden dataset without uploading it. Duplicate groups with row and line numbers, an identity key you choose, and the sample size left after a dedupe.
- Calculator
Near-Duplicate Row Detector: Find Eval Rows That Differ Only in Formatting
Find rows in an eval set that are the same test written twice. Normalize the columns you pick, group rows whose normalized text matches, and read the exact transform, the raw variants, and a CSV.
- Calculator
Cross-Split Overlap Checker: Find Eval Items in Both Train and Test
Compare two eval splits in your browser and see which items sit on both sides, with a separate overlap rate against the row count of each split, exact and normalized matching, and the lines to open.
- Calculator
N-Gram Overlap Checker: Find Contamination Between Two Eval Files
Check a test set against a training file for shared word sequences. Pick n, read the per-row dirty-token share, see the matching text in both files, and export a CSV. Both files stay in your browser.
- Calculator
Slice and Class Balance Analyzer: See What an Eval Set Contains
Read what an eval set holds before you quote a pass rate. Count rows per slice, per class, and per control flag, put an interval on every share, and see the smallest slice named.
- Calculator
Evaluation Card Generator: Markdown and JSON on a Published Schema
Write one evaluation down so a second person can read the number correctly, or run it again. Fill the form, export a Markdown card and a JSON card that validates against a schema published here.
- Calculator
Independent Two-Proportion Test Calculator for Eval Arms
Compare two pass rates measured on different case sets: a pooled two-proportion z test, Fisher exact when an arm is small, and an interval on the difference.
- Calculator
Evaluation Coverage Matrix: Which Cells Have No Cases
Put what you test on one axis and what you test it across on the other. The grid shows which cells have cases, which are empty, which are too small to carry an interval, and what each gap costs you.
- Calculator
Evaluator and Metric Navigator: Which Grader, Which Number
Answer four questions about oracle type, grader, system count, and failure mode. Get the admissible evaluator, inadmissible ones and why, the metric family, and a route to the tool that runs it.
- Calculator
What Does This Evaluation Prove?
Select the evidence your evaluation produced, see which claims it supports, which it cannot, and the inference error behind each gap. Each unsupported claim names the instrument that would close it.
- Calculator
Golden Set Blueprint Builder: Seed a Test Set From a Task Family
Choose a task family, load its case taxonomy and starter rows, then edit, add and export. Every row carries an oracle type and a provenance tag, and the export validates against the eval-set schema.
- Calculator
Claim-Evidence Matrix: What Backs Each Deployment Claim
One row per deployment claim, eleven evidence dimensions from relevance and representativeness to independence and applicability, with named rate-down reasons and residual deficits.
- Calculator
Reproducibility Checklist and Manifest Validator
Check whether your write-up records what someone else would need to rerun your eval: a checklist against the published card schema, and a validator for a pasted or uploaded manifest.