INSTRUMENT | Eval statistics
Claim-Evidence Matrix: What Backs Each Deployment Claim
One row per deployment claim, eleven evidence dimensions from relevance and representativeness to independence and applicability, with named rate-down reasons and residual deficits.
Each row is a deployment claim. Behind it: the artifacts, and a per-dimension status across eleven evidence dimensions. Where a dimension is limited, a named reason says why. Residual deficits are listed, not scored. Built on the evidence quality framework and the methodology in reporting LLM-as-a-judge evaluations. No composite score, no certainty grade, no sufficiency verdict. To check what a single evaluation run proves, use what does this evaluation prove.
Assess one claim
Evidence register
No claims registered yet. Use the form above to assess your first claim.
Adequate Limited Not recorded Not applicable
Worked example (static preview). With JavaScript, the register is interactive.
| Claim | Class | Rel | Repr | Meth | Stat | Cons | Repro | Trac | Ind | Rec | Comp | App | Deficits | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Intent routing pass rate on the September sample is 91% (95% interval 86 to 95, n = 180) | Capability on this set | + | ~ | + | ~ | + | ? | + | + | + | ~ | + | Repr: convenience sample, not production traffic; Stat: n = 180, wide interval on tail slice; Comp: refund path not tested |
Gaps and deficits
No gaps to show.
Eleven evidence dimensions
Each claim is assessed against eleven dimensions defined in the evidence quality framework. A dimension marked "limited" carries a written reason; "not recorded" flags a gap the reader has not yet inspected; "not applicable" means the dimension does not bear on this claim.
- Relevance
-
Does the evidence address this claim in this context of use?
Rate down when: Wrong population, wrong comparator, or wrong outcome for the claim.
- Representativeness
-
Was the sample drawn from the population the claim covers, including its tail?
Rate down when: Sample drawn from a narrower population than the claim addresses.
- Methodological soundness
-
Was the evaluator valid, were slices declared in advance, was grading auditable?
Rate down when: Evaluator not validated, post-hoc slicing, or unauditable scoring.
- Statistical adequacy
-
Is n sufficient for the precision claimed, per slice, with intervals reported?
Rate down when: Sample too small for the claimed precision, or intervals not reported.
- Consistency
-
Do independent runs, slices, and evaluators agree?
Rate down when: Runs, slices, or evaluators disagree beyond the reported interval.
- Reproducibility
-
Could someone else re-run it with versions, prompts, seeds, dates, and harness pinned?
Rate down when: Missing version pins, unpublished prompts, or undisclosed seeds.
- Traceability
-
Does each number trace to specific cases and a specific system version?
Rate down when: Aggregate without case-level traceability or unspecified system version.
- Independence
-
Who produced the evidence, and did they have an interest in the result?
Rate down when: Evidence produced by a party with a stake in the outcome.
- Recency
-
Does the evidence describe the currently deployed configuration?
Rate down when: Evidence predates the deployed version or a material configuration change.
- Completeness
-
What was not tested, stated rather than omitted?
Rate down when: Known gaps not stated, or selective reporting of favorable slices.
- Applicability
-
Does the deployment context match the evaluation context?
Rate down when: Evaluation context differs materially from the deployment context.
Admissible artifacts include a judge calibration log and a judge version card. The evidence pack feeds the release gate designer.
Export
Claim-evidence register
Deployment evidence dossier
Import
A previously exported register file. Nothing leaves your browser.
JSON. Drag one here, or use the box below.
Check this value.
Common questions
Why no composite score?
The eleven dimensions are not commensurable: good reproducibility does not offset an invalid evaluator. The evidence quality framework gives the full argument against a composite; the short version is that any weighting is a value judgment disguised as arithmetic, and a composite is a target that will be optimized.
What does "not recorded" mean versus "not applicable"?
"Not recorded" means you have not yet assessed that dimension for this claim. It is a visible gap, flagged in the deficits panel. "Not applicable" means the dimension genuinely does not bear on this claim and is not counted as a gap. The two are never interchangeable: a dimension you skipped because you ran out of time is "not recorded," not "not applicable."
What is a residual deficit?
A deficit that remains after you have assessed every dimension. It is a named doubt in the evidence base that you choose to carry rather than resolve, stated explicitly so a reader knows it exists. An assurance case without stated residual deficits is not making a stronger argument; it is making an incomplete one.