LatentEval

INSTRUMENT | Reliability testing

Evaluation Blueprint Builder: Write the Argument Before You Run the Eval

State each claim your evaluation will make, name the evidence that would support or refute it, and write the inference step connecting the two, all before any results exist.

Write the argument an evaluation will make before results exist, built on the measurement discipline described in how to measure agent reliability. Each claim names the evidence that would support or refute it and the inference step connecting the two: the reason the evidence is relevant, not just the evidence itself. Import a context-of-use profile to carry its design consequences forward. Check slice balance with the class balance analyzer and estimate execution cost with the eval budget calculator. To preregister the plan, freeze it, and log results once the eval runs, use the eval run register.

System

Check this value.

Import a context-of-use profile

Paste or pick a JSON export from the context-of-use profiler. Nothing is uploaded. The profile fills the system name and pre-fills the slices list and the threats list from its consequences; the claims, evaluators and run plan you write yourself.

JSON. Drag one here, or use the box below.

Check this value.

Claims

Each claim is something this evaluation would establish if the evidence comes back as expected. The inference step says WHY the evidence is relevant: the reasoning that connects a measurement to a conclusion. The defeater names what could make the inference fail even if the evidence looks right.

Claim C1

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Evidence checklist

Mark each element the evidence above includes. The gap check reports required elements not marked "yes."

Population

Check this value.

Check this value.

Slices

Named subsets of the population that the evaluation must report on separately, with their minimum case counts. Use the golden-set size planner to choose a minimum that supports the claim you need.

Slice 1

Check this value.

Check this value.

Check this value.

Evaluators

Who or what grades each case. Name the evaluator and say what kind it is: a programmatic check, a human rater, an LLM judge with agreement evidence, a domain expert.

Evaluator 1

Check this value.

Check this value.

Threats to validity

Anything that could make the evaluation's conclusions wrong even if the measurements are correct. Name the threat and say what you will do about it.

Threat 1

Check this value.

Check this value.

Run plan

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Notes

Check this value.

Claims

Population
not described
Slices
0
Evaluators
0
Threats recorded
0
Run plan
not present
Imported profile
none
Export the blueprint

Markdown is the document for a design review. JSON is the same document and loads back into this form.

How this is calculated

This page writes down the argument, not the result. Nothing is scored, graded, rated or indexed, and there is no headline number. The output is a structured record of what each claim rests on, with two diagnostic lists that follow from it.

Evidence gaps list the evidence elements a claim class requires that the reader has not yet confirmed. A claim class is a category of conclusion ("the system passed this share of these cases") and each class requires specific evidence elements (a case set with a stated population, an interval on every rate). The checklist under each claim lets you mark each element as present, absent, or not yet checked.

Unresolved items are fields the blueprint expects and does not have yet: a claim with no inference step, a population with no description, a run plan with no stopping rule. Each one is named so you can close it or decide it does not apply.

A plan's cost section draws on the cost per successful task and the eval budget calculator. The rows the plan must cover come from the agent action risk matrix and the irreversible action inventory, and its suites become release conditions in the release gate designer.

Questions

Questions

What is the inference step?

The reason the evidence you named is relevant to the claim you made. In the GSN (Goal Structuring Notation) family, this is the strategy node: the step between a goal and the evidence that discharges it. Most evaluation writeups leave it implicit, which is where the argument breaks down. Two people can read the same evidence and disagree about what it establishes, because neither wrote down the reasoning.

What is a defeater?

Something that would make the inference fail even if the evidence looks right. A defeater for "the pass rate is above 90 percent" might be "the cases were seen during development, so the rate measures tuning, not capability." Not every claim has a known defeater, but writing one down is the discipline that keeps the argument honest.

Why no score or readiness level?

Because an evaluation blueprint is a plan, not a result. There is nothing to score yet. The output of this page is a record of what you will measure, why you will measure it that way, and what could go wrong, enough for a reviewer to say whether the argument holds before any data is collected.

Where does the context-of-use profile come from?

The context-of-use profiler. Export a JSON from that page, then pick or paste it into the import section here. The profile fills the system name and pre-fills the slices list and the threats list from its consequences; the claims, evaluators and run plan you write yourself.