INSTRUMENT | Reliability testing
Evaluation Blueprint Builder: Write the Argument Before You Run the Eval
State each claim your evaluation will make, name the evidence that would support or refute it, and write the inference step connecting the two, all before any results exist.
Write the argument an evaluation will make before results exist, built on the measurement discipline described in how to measure agent reliability. Each claim names the evidence that would support or refute it and the inference step connecting the two: the reason the evidence is relevant, not just the evidence itself. Import a context-of-use profile to carry its design consequences forward. Check slice balance with the class balance analyzer and estimate execution cost with the eval budget calculator. To preregister the plan, freeze it, and log results once the eval runs, use the eval run register.
Showing your last valid result. Update the inputs above to recompute.
Claims
Evidence gaps
Each row is an evidence element the claim class requires that has not been marked "yes" in the checklist. A gap here means the argument has a step without the artifact it rests on.
| Claim | Claim class | Missing element |
|---|
- Population
- not described
- Slices
- 0
- Evaluators
- 0
- Threats recorded
- 0
- Run plan
- not present
- Imported profile
- none
Markdown is the document for a design review. JSON is the same document and loads back into this form.
How this is calculated
This page writes down the argument, not the result. Nothing is scored, graded, rated or indexed, and there is no headline number. The output is a structured record of what each claim rests on, with two diagnostic lists that follow from it.
Evidence gaps list the evidence elements a claim class requires that the reader has not yet confirmed. A claim class is a category of conclusion ("the system passed this share of these cases") and each class requires specific evidence elements (a case set with a stated population, an interval on every rate). The checklist under each claim lets you mark each element as present, absent, or not yet checked.
Unresolved items are fields the blueprint expects and does not have yet: a claim with no inference step, a population with no description, a run plan with no stopping rule. Each one is named so you can close it or decide it does not apply.
A plan's cost section draws on the cost per successful task and the eval budget calculator. The rows the plan must cover come from the agent action risk matrix and the irreversible action inventory, and its suites become release conditions in the release gate designer.
Questions
Questions
What is the inference step?
The reason the evidence you named is relevant to the claim you made. In the GSN (Goal Structuring Notation) family, this is the strategy node: the step between a goal and the evidence that discharges it. Most evaluation writeups leave it implicit, which is where the argument breaks down. Two people can read the same evidence and disagree about what it establishes, because neither wrote down the reasoning.
What is a defeater?
Something that would make the inference fail even if the evidence looks right. A defeater for "the pass rate is above 90 percent" might be "the cases were seen during development, so the rate measures tuning, not capability." Not every claim has a known defeater, but writing one down is the discipline that keeps the argument honest.
Why no score or readiness level?
Because an evaluation blueprint is a plan, not a result. There is nothing to score yet. The output of this page is a record of what you will measure, why you will measure it that way, and what could go wrong, enough for a reviewer to say whether the argument holds before any data is collected.
Where does the context-of-use profile come from?
The context-of-use profiler. Export a JSON from that page, then pick or paste it into the import section here. The profile fills the system name and pre-fills the slices list and the threats list from its consequences; the claims, evaluators and run plan you write yourself.