LatentEval

INSTRUMENT | Eval statistics

Claim-Evidence Matrix: What Backs Each Deployment Claim

One row per deployment claim, eleven evidence dimensions from relevance and representativeness to independence and applicability, with named rate-down reasons and residual deficits.

Each row is a deployment claim. Behind it: the artifacts, and a per-dimension status across eleven evidence dimensions. Where a dimension is limited, a named reason says why. Residual deficits are listed, not scored. Built on the evidence quality framework and the methodology in reporting LLM-as-a-judge evaluations. No composite score, no certainty grade, no sufficiency verdict. To check what a single evaluation run proves, use what does this evaluation prove.

Assess one claim

The claim

Check this value.

Narrows which evidence elements matter most.

Check this value.

Check this value.

Check this value.

Evidence dimensions

Does the evidence address this claim in this context of use?

Check this value.

Was the sample drawn from the population the claim covers, including its tail?

Check this value.

Was the evaluator valid, were slices declared in advance, was grading auditable?

Check this value.

Is n sufficient for the precision claimed, per slice, with intervals reported?

Check this value.

Do independent runs, slices, and evaluators agree?

Check this value.

Could someone else re-run it with versions, prompts, seeds, dates, and harness pinned?

Check this value.

Does each number trace to specific cases and a specific system version?

Check this value.

Who produced the evidence, and did they have an interest in the result?

Check this value.

Does the evidence describe the currently deployed configuration?

Check this value.

What was not tested, stated rather than omitted?

Check this value.

Does the deployment context match the evaluation context?

Check this value.

Evidence register

Adequate Limited Not recorded Not applicable

Worked example (static preview). With JavaScript, the register is interactive.

Claim Class Rel Repr Meth Stat Cons Repro Trac Ind Rec Comp App Deficits
Intent routing pass rate on the September sample is 91% (95% interval 86 to 95, n = 180) Capability on this set + ~ + ~ + ? + + + ~ + Repr: convenience sample, not production traffic; Stat: n = 180, wide interval on tail slice; Comp: refund path not tested

Gaps and deficits

No gaps to show.

Eleven evidence dimensions

Each claim is assessed against eleven dimensions defined in the evidence quality framework. A dimension marked "limited" carries a written reason; "not recorded" flags a gap the reader has not yet inspected; "not applicable" means the dimension does not bear on this claim.

Relevance

Does the evidence address this claim in this context of use?

Rate down when: Wrong population, wrong comparator, or wrong outcome for the claim.

Representativeness

Was the sample drawn from the population the claim covers, including its tail?

Rate down when: Sample drawn from a narrower population than the claim addresses.

Methodological soundness

Was the evaluator valid, were slices declared in advance, was grading auditable?

Rate down when: Evaluator not validated, post-hoc slicing, or unauditable scoring.

Statistical adequacy

Is n sufficient for the precision claimed, per slice, with intervals reported?

Rate down when: Sample too small for the claimed precision, or intervals not reported.

Consistency

Do independent runs, slices, and evaluators agree?

Rate down when: Runs, slices, or evaluators disagree beyond the reported interval.

Reproducibility

Could someone else re-run it with versions, prompts, seeds, dates, and harness pinned?

Rate down when: Missing version pins, unpublished prompts, or undisclosed seeds.

Traceability

Does each number trace to specific cases and a specific system version?

Rate down when: Aggregate without case-level traceability or unspecified system version.

Independence

Who produced the evidence, and did they have an interest in the result?

Rate down when: Evidence produced by a party with a stake in the outcome.

Recency

Does the evidence describe the currently deployed configuration?

Rate down when: Evidence predates the deployed version or a material configuration change.

Completeness

What was not tested, stated rather than omitted?

Rate down when: Known gaps not stated, or selective reporting of favorable slices.

Applicability

Does the deployment context match the evaluation context?

Rate down when: Evaluation context differs materially from the deployment context.

Admissible artifacts include a judge calibration log and a judge version card. The evidence pack feeds the release gate designer.

Export

Claim-evidence register

Export register

Deployment evidence dossier

Export dossier

Import

A previously exported register file. Nothing leaves your browser.

JSON. Drag one here, or use the box below.

Check this value.

Common questions

Why no composite score?

The eleven dimensions are not commensurable: good reproducibility does not offset an invalid evaluator. The evidence quality framework gives the full argument against a composite; the short version is that any weighting is a value judgment disguised as arithmetic, and a composite is a target that will be optimized.

What does "not recorded" mean versus "not applicable"?

"Not recorded" means you have not yet assessed that dimension for this claim. It is a visible gap, flagged in the deficits panel. "Not applicable" means the dimension genuinely does not bear on this claim and is not counted as a gap. The two are never interchangeable: a dimension you skipped because you ran out of time is "not recorded," not "not applicable."

What is a residual deficit?

A deficit that remains after you have assessed every dimension. It is a named doubt in the evidence base that you choose to carry rather than resolve, stated explicitly so a reader knows it exists. An assurance case without stated residual deficits is not making a stronger argument; it is making an incomplete one.