LatentEval

INSTRUMENT | Eval statistics

What Does This Evaluation Prove?

Select the evidence your evaluation produced, see which claims it supports, which it cannot, and the inference error behind each gap. Each unsupported claim names the instrument that would close it.

Tick what your evaluation actually produced and see which claims that evidence supports. Eleven claim classes, each with a named inference error. Built on the significance walkthrough and the fourteen evidence elements below. This page judges one evaluation run against the claim it is being asked to carry; to track many claims and the artifacts behind them across a deployment, use the claim-evidence matrix. A known cross-split overlap is a limitation the evidence carries. For background on what each eval type can certify, see what LLM evals are.

What did this evaluation produce?
Claims not supported

What is missing for each claim

ClaimWhat it saysWhat is missingWhat would establish it
Capability on this set The system passed this share of these cases, measured at this date. A set of cases with a written statement of the population they were drawn from.; A confidence interval reported beside every rate.; The exact system version and the date the measurement was taken.; A check that the cases did not reach the system during training or development. Pass rate ci calculator, Golden set size planner
Reliability across runs The system does this consistently, not once. A set of cases with a written statement of the population they were drawn from.; The same case run more than once, with the results kept apart.; A confidence interval reported beside every rate. Reliability at k estimator, Repeated run variance planner
Improvement over baseline This version is better than the one it replaces, on these cases. A set of cases with a written statement of the population they were drawn from.; The previous version measured on the same cases, or on an independent sample with the difference reported with its own interval.; A confidence interval reported beside every rate.; One thing different between the two arms, with everything else held.; A check that the cases did not reach the system during training or development. Mcnemar test calculator, Eval ab comparator, Paired vs independent design chooser
No regression on held cases Nothing that used to work stopped working. A set that was not looked at while the system was being built or tuned.; The previous version measured on the same cases, or on an independent sample with the difference reported with its own interval.; A confidence interval reported beside every rate.; The exact system version and the date the measurement was taken.; A check that the cases did not reach the system during training or development. Cross split overlap checker, Golden set version tracker, Eval run register, Mcnemar test calculator
Rare severe failure bound The severe-failure rate is below the upper limit of the interval computed at the n you counted for that class. The severe failures counted as their own class, with the severe class's own n stated.; A set of cases with a written statement of the population they were drawn from.; A confidence interval reported beside every rate. Pass rate ci calculator, Golden set size planner, Acceptable error rate calculator, Acceptance criteria builder
Per-slice parity The system works about as well on this subset as on the whole. Named slices, with how many cases fell into each one.; A confidence interval reported beside every rate. Slice and class balance analyzer, Golden set size planner
Robustness to stated perturbation The system holds up under this named change to the input. Cases altered by a named, repeatable transformation, kept beside their originals.; The same system measured on the unaltered originals, beside the perturbed cases.; A confidence interval reported beside every rate. Prompt robustness analyzer, Mcnemar test calculator
Evaluator validity The grader that produced these numbers agrees with people on these items. Agreement against human labels on your own items, reported as a statistic with its interval.; The size of the labeled sample stated. Inter rater reliability calculator, Judge validation report builder, Judge agreement tracker
Capability in deployment The system will work this well for the people who actually use it. Cannot be established offline Eval run register, Review threshold optimizer
Robustness to unseen shift The system will hold up when the inputs change in ways nobody wrote down. Cannot be established offline Prompt robustness analyzer, Golden set version tracker
Causal attribution of a change This specific change is what moved the number. One thing different between the two arms, with everything else held.; The previous version measured on the same cases, or on an independent sample with the difference reported with its own interval.; The same case run more than once, with the results kept apart.; A confidence interval reported beside every rate. Change impact worksheet, Eval ab comparator, Repeated run variance planner
Named inference errors

The inference error behind each unsupported claim

ClaimThe inference errorActive
Capability on this set A rate read without its interval. A pass rate quoted to one decimal place on 40 cases is reporting noise as precision. Active
Reliability across runs pass@k read as reliability. pass@k rises with k by construction; a run-once product experiences pass@1, and a must-succeed-every-time product experiences pass^k. Active
Improvement over baseline Two rates compared with no interval on the difference. Pairing on the same cases removes case-difficulty variance; an independent design needs the two-proportion interval instead. Active
No regression on held cases A regression claim with no prior measurement. Nothing establishes what used to work unless the previous version was measured on the same cases. Active
Rare severe failure bound "Zero failures" read as "no failures happen". Zero in n bounds the rate at about 3/n (one-sided 95 percent); 0 of 100 is consistent with a true rate up to 3 percent. Active
Per-slice parity The average read as the slice. A 97 percent pass rate is compatible with total failure on a severe slice that is 3 percent of the set, because the slice is a rounding error in the mean. Active
Robustness to stated perturbation A perturbation set with no original beside it. Without the unperturbed cases the drop cannot be separated from the difficulty of the new cases. Active
Evaluator validity Judge agreement read as correctness. A judge aligned to human preference can be near chance on objectively correct versus subtly incorrect pairs. Active
Capability in deployment Benchmark rank read as task fitness. Always active
Robustness to unseen shift Repeated runs read as robustness. Re-running the same cases reduces uncertainty about run-to-run variance, not about population coverage. Always active
Causal attribution of a change Two things changed at once and the credit given to the one you were interested in. Active

An active error means the claim is unsupported and a reader drawing the claim from this evaluation is making this error. An always-active error belongs to a claim that no offline evaluation can establish, whatever evidence is ticked. A not-active error means the claim is supported.

How this is calculated

Pure set containment. For each of eleven claim classes, the tool checks whether the evidence elements you ticked include everything the claim requires. Two claim classes can never be established by an offline evaluation, whatever is ticked: they appear in the unsupported list with a separate explanation.

Evidence elements. Fourteen items form the entire input surface. Each is a concrete thing an evaluation either produced or did not: a case set with a stated population, an interval on every rate, repeated runs per case, a prior version measured on the same cases, unperturbed originals measured beside the perturbed set, a held-out set, an evaluator with agreement evidence, a labeled sample size, a severe-failure class with its own n, a contamination check, a pinned version, a perturbation set, one changed variable, and declared slices with per-slice counts.

Claim classes. Each claim has a statement (what you can say), a requires list (the evidence it needs), a common error (the inference mistake a reader makes when the evidence is missing), and a closes-with list (tools on this site that run the missing number).

No composite score. The output is a list of supported and unsupported claims. There is no readiness score, no maturity band, and no count-over-total. The value is in which claims stand and which do not, not in how many.

Judge calibration sits in the judge agreement tracker; the fields this page reads are the same ones the evaluation card generator publishes.

Questions

Questions

What does "always active" mean in the inference errors table?

An always-active error belongs to a claim that no offline evaluation can establish, whatever evidence you tick. The error is inherent to the gap between an offline measurement and the deployment claim.

Why are two claims never supported?

Capability in deployment and robustness to unseen shift require evidence from the running system, not from an offline evaluation. No combination of the fourteen evidence elements can bridge that gap.

Can I use this tool for a multi-run evaluation?

Yes. Tick "repeated runs per case" if you ran the system more than once on each case. The tool checks whether your evidence set supports claims about reliability across runs.

What is the difference between this tool and the claim-evidence matrix?

This page judges one evaluation run against the claim it is being asked to carry. The claim-evidence matrix tracks many claims, the artifacts behind them, and per-dimension evidence quality across a deployment.

Why is there no score?

The output is which claims your evidence supports and which it does not. A count of supported claims is a target that invites collecting the cheap evidence, not the evidence the claim actually needs. The value is in which claims stand, not in how many.