LatentEval

For builders

Evidence quality for AI claims: eleven dimensions, no score

Eleven dimensions for appraising evaluation evidence, named rate-down reasons for each, and a five-part argument against collapsing them into a composite number.

For builders

In brief

4 POINTS
  • Certainty belongs to a specific claim at a specific threshold, not to a study design or a dossier.
  • Evidence classes are contexts, not a ladder: a single hierarchy level spreads across every certainty level once you re-appraise it claim by claim.
  • A composite evidence score becomes a target, and weighting eleven incommensurable dimensions is a value judgment dressed as arithmetic.
  • Nearly all AI evaluation evidence is self-produced, and that is worth stating directly rather than burying in a dimension score.

You have evaluation results. Numbers in a table, maybe confidence intervals, maybe a pass rate on a golden set. The question you actually need answered is not “how good are these numbers” but “how much should this claim, backed by this evidence, change what I decide to do?” That second question turns on the quality of the evidence behind each specific claim, and it does not have a single-number answer.

This framework is the method behind the claim-evidence matrix. It rests on our work in reporting LLM-as-a-judge evaluations and the four ways an eval number lies, and draws its structure from the GRADE Working Group’s 2017 clarification of certainty of evidence (Hultcrantz et al.).

Certainty is claim-specific

The first correction this framework makes is structural. Certainty of evidence is not a property of a study, a dataset, or a dossier. It is a property of a specific claim at a specific decision threshold. The GRADE Working Group’s own 2017 clarification states this directly: certainty is confidence that the true effect lies on one side of a specified threshold or within a chosen range. Two consequences follow.

First, the same evaluation run can support one claim with high certainty and another with low certainty. A benchmark that measures accuracy on English-language inputs supports the claim “the system handles English queries” far more strongly than the claim “the system handles multilingual queries.” The evidence is identical; the claims are not.

Second, study design sets a starting point, not a verdict. A rigorous observational body of evidence can be rated up to moderate or high on GRADE’s rate-up criteria: a large effect, a dose-response gradient, or confounding that would work against the observed effect. A randomized controlled trial can fall to low certainty if the sample is too small for the precision claimed or if the outcome measured is not the outcome that matters. Design is a starting position, not a floor.

The eleven dimensions

Each claim is assessed against eleven dimensions. These are compiled from GRADE’s rate-down domains, assurance-deficit categories, audit practice, and procurement guidance. They are our compilation and labeled as such.

Relevance. Does the evidence address this claim in this context of use? An evaluator that measures the wrong construct, tests the wrong population, or compares against the wrong baseline is indirect evidence regardless of its rigor. GRADE calls this “indirectness.”

Representativeness. Was the sample drawn from the population the claim covers, including its tail? A benchmark that excludes edge cases, rare languages, or adversarial inputs is representative of the easy middle, not of the deployment population.

Methodological soundness. Was the evaluator valid, were slices declared in advance, was the grading auditable? GRADE calls this “risk of bias.” In AI evaluation, the most common methodological failure is using an unvalidated evaluator: a judge LLM whose agreement with human raters has never been measured for this task.

Statistical adequacy. Is n sufficient for the precision claimed, per slice, with intervals reported? A pass rate of 94% on 50 cases has a 95% Clopper-Pearson interval of roughly 83% to 99% (Wilson: 84% to 98%). Reporting the point estimate without the interval is reporting the least informative number the evaluation produced. GRADE calls this “imprecision.” The golden set size planner sizes a set for a stated baseline rate and the drop worth catching; whether that n is adequate is settled by the claim, not by the planner.

Consistency. Do independent runs, slices, and evaluators agree? If three runs on the same data produce accuracy figures of 87%, 92%, and 78%, the claim “accuracy is above 90%” rests on one of three runs. GRADE calls this “inconsistency.” Plan the trial count with the repeated run variance planner.

Reproducibility. Could someone else re-run this evaluation? That requires versions, prompts, seeds, dates, and harness all pinned and disclosed. Without them, the number is a report about a moment, not a measurement someone can verify. The reproducibility checklist validator checks whether the pins are in place.

Traceability. Does each number trace to specific cases and a specific system version? An aggregate without case-level provenance hides which cases the system failed on and whether those cases matter for the claim at hand. The evaluation card generator produces a card that records this lineage.

Independence. Who produced the evidence, and did they have an interest in the result? This dimension is worth stating honestly: almost all AI evaluation evidence today is self-produced. First-party evidence is not automatically wrong, but it is never independent. A framework that scores independence the same way it scores statistical adequacy will give nearly every evaluation a failing mark on one dimension and create an incentive to invent independence rather than acknowledge its absence. The judge validation report builder helps separate evaluator validity from the team that built the system.

Recency. Does the evidence describe the currently deployed configuration? A model that was fine-tuned, re-prompted, or updated since the evaluation was run has evidence that describes a different system. How different depends on what changed, but the evidence is no longer about the thing making decisions in production.

Completeness. What was not tested, stated rather than omitted? GRADE’s analogue is publication bias: selective reporting of favorable results. In AI evaluation, the most common form is unstated scope. The evaluation tested English; the deployment serves twelve languages. The gap is not in the evidence that exists but in the evidence that was never produced. The eval dataset schema validator checks file-level fields, and its duplicate-id and completeness section counts empty cells.

Applicability. Does the deployment context match the evaluation context? An evaluation run in a sandbox with synthetic data tells you what the system does under those conditions. Whether it does the same thing with real user inputs, real latency constraints, and real upstream data is a separate question.

Evidence classes are contexts, not a ladder

The traditional “levels of evidence” pyramid ranks study designs in a general-purpose hierarchy: systematic reviews at the top, expert opinion at the bottom. The critique, well-supported in the GRADE literature, is that this ranking does not survive contact with specific claims. Two bodies of evidence at the same hierarchy level can end up at very different certainties once each is appraised against a specific claim and threshold.

For AI evaluation evidence, the relevant classes are not study designs but evidence contexts: a controlled benchmark, a production trace analysis, a red-team exercise, a user-reported incident log, a third-party audit, a vendor-supplied evaluation card. None of these is inherently above or below any other. A controlled benchmark on a representative sample with a validated evaluator can support a narrow claim with high certainty. A production trace analysis covering six months of real traffic can support a deployment-scope claim that no benchmark can reach. The question is which dimensions each class covers well and which it leaves exposed.

The claim-evidence matrix does not ask “what level is this evidence?” It asks, for each dimension, “is this dimension adequate, limited, not recorded, or not applicable for this claim?” The answer depends on the claim, the evidence, and the context, not on a fixed hierarchy.

The five-part argument against a composite score

State this plainly, because declining to score is a design decision that needs its own evidence.

One. Certainty is claim-specific. A composite score over a dossier is a score for no particular claim and therefore supports no particular decision. If someone asks “what is the overall evidence quality?” the honest answer is that the question is not well-formed until it names a claim and a threshold.

Two. The eleven dimensions are not commensurable. Excellent reproducibility does not compensate for an invalid evaluator. High statistical power does not offset the fact that the sample was drawn from the wrong population. Averaging across dimensions implies a trade-off rate that nobody can justify and nobody has measured.

Three. A composite score is a target, and it will be optimized. An organization that receives a “B+” on evidence quality will work to reach “A” by improving whichever dimension is cheapest to improve, not whichever dimension matters most for the claim at hand. The optimization is rational and produces the wrong outcome.

Four. Any weighting is a value judgment presented as arithmetic. A framework that weights independence at 15% and recency at 10% has not measured the relative importance of those dimensions. It has made a policy choice, embedded it in a formula, and removed it from the conversation where it belongs.

Five. The useful output is per-claim, per-dimension statuses with named rate-down reasons plus the explicit list of residual deficits. That is what a confidence argument produces and what a buyer or a governance board can act on. GRADE arrives at a certainty level because it has a comparable body of studies to pool; a single bespoke evaluation does not, so this framework borrows the named-reason structure and stops short of the level. It is less convenient than a letter grade and more honest about what the evidence says.

Rate-down reasons

When a dimension is marked “limited,” the framework requires a written reason. This is not a documentation burden for its own sake. The reason is the finding. “Independence: limited because all evidence is first-party” tells a reader something actionable. “Independence: 2 out of 5” tells them nothing they can act on except to raise the number.

Rate-down reasons borrow GRADE’s named-reason structure: risk of bias, inconsistency, indirectness, imprecision, and publication bias as starting points, extended with the AI-specific concerns (evaluator validity, version pinning, deployment context mismatch) that the eleven dimensions capture. Each rate-down reason is a specific, inspectable claim about a specific gap, not a position on a scale.

Residual deficits

An assurance case that carries no stated uncertainties is not making a stronger argument. It is making an incomplete one. The safety-assurance literature treats the gap between “we have evidence” and “the evidence supports the claim” as a first-class object: an assurance deficit is any identified doubt in a claim, an inference step, or an item of evidence.

The claim-evidence matrix carries a residual-deficits field on every claim row. These are the doubts you chose to carry rather than resolve. Stating them explicitly means a reader knows they exist, a governance board can weigh them, and a future evaluation can target them. Omitting them means the same doubts exist but nobody downstream knows about them.

A compelling confidence argument identifies its deficits and says why the residual ones are acceptable. That “why” is a judgment the people making the deployment decision need to own, not a judgment a framework should hide inside a formula.

Independence deserves honesty

Most AI evaluation evidence is produced by the team that built the system. That is worth saying directly rather than hiding behind a dimension that almost every evaluation fails.

Independence is not a binary. A first-party evaluation with a pre-registered protocol, a disclosed evaluator, and case-level traceability is different from a first-party evaluation with none of those things. But neither is independent in the sense that an external audit or a third-party red team is independent. The framework marks the dimension, states the reason, and leaves the judgment to the reader. It does not pretend that internal testing is independent because the testing team reported to a different manager.

The practical consequence is that a deployment evidence dossier will almost always carry at least one “limited” dimension on independence. That is the honest state of the field, not a defect in the framework.

How the claim-evidence matrix operationalizes this

The claim-evidence matrix is the instrument that turns this methodology into a working register. Each row is a deployment claim. Behind it: the artifacts that back it (evaluation cards, run cards, validation reports, manifests, datasets), and a per-dimension status on each of the eleven dimensions described here.

Where a dimension is adequate, the row records it and moves on. Where it is limited, the row carries a written reason. Where it has not been recorded, the gap is visible rather than absent. Where the dimension does not apply to this claim, it says so. Residual deficits are listed explicitly. The whole register exports as a JSON file for programmatic use, and the dossier exports as a Markdown document for human review.

No row produces a composite score. No view aggregates dimensions into a grade. No export calculates a percentage of adequate dimensions. The product is the named state of each dimension on each claim, and the residual deficits the organization chose to carry.

What this framework does not do

It does not rank study designs. It does not produce a letter grade, a readiness level, or a maturity band. It does not tell you whether your evidence is “enough” in the abstract, because that question requires a claim, a threshold, and a decision context that the framework does not choose for you.

It does tell you, per claim, which dimensions hold up, which ones fall short and why, and what remains uncertain. That is a less satisfying answer than a score, and it is the one the evidence actually supports. GRADE was built for health-outcome interventions with a literature of comparable studies; AI deployment evidence is usually a single bespoke evaluation with no comparator body. The framework borrows GRADE’s structure, not its machinery: per-claim assessment, rate-down for named inspectable reasons, threshold-relative rather than design-ranked. That structure transfers. The claim classes, the evidence kinds, and the dimension definitions are ours, compiled from the literature and from the failure modes we have seen in practice.

The tools that operationalize these dimensions are collected in the evidence and reproducibility topic.