INSTRUMENT | Eval statistics
Golden Set Blueprint Builder: Seed a Test Set From a Task Family
Choose a task family, load its case taxonomy and starter rows, then edit, add and export. Every row carries an oracle type and a provenance tag, and the export validates against the eval-set schema.
Start from a task-family blueprint: pick a family, load its case families, slice families and starter rows. Edit them in the table below, add your own rows, and export a dataset that validates against the published eval-set schema with a case profile on every row. The oracle type on each case gates which evaluator can grade it. The structure draws on our research into how agent errors go undetected and what an agent evaluation has to measure. Each task-family pack has its own guide with the reasoning, failure analysis and evaluator guidance behind its rows; this page emits the rows themselves.
Showing your last valid result. Update the inputs above to recompute.
No rows yet.
Validation
10 rows, 10 valid.
- Row 1: case.starter-unvalidated - provenance is starter-example, so this case is starting material, not a validated answer.
- Row 2: case.starter-unvalidated - provenance is starter-example, so this case is starting material, not a validated answer.
- Row 3: case.starter-unvalidated - provenance is starter-example, so this case is starting material, not a validated answer.
- Row 4: case.starter-unvalidated - provenance is starter-example, so this case is starting material, not a validated answer.
- Row 5: case.starter-unvalidated - provenance is starter-example, so this case is starting material, not a validated answer.
- Row 6: case.starter-unvalidated - provenance is starter-example, so this case is starting material, not a validated answer.
- Row 7: case.starter-unvalidated - provenance is starter-example, so this case is starting material, not a validated answer.
- Row 8: case.starter-unvalidated - provenance is starter-example, so this case is starting material, not a validated answer.
- Row 9: case.starter-unvalidated - provenance is starter-example, so this case is starting material, not a validated answer.
- Row 10: case.starter-unvalidated - provenance is starter-example, so this case is starting material, not a validated answer.
Export
Each format writes the same rows. Canonical CSV, JSON and JSONL are the eval-set schema itself. The vendor formats reshape the same data for that vendor's harness; the Lossy column below names any field the format cannot carry.
How this works
Every evaluation case declares an oracle type: the form its expected answer takes. The oracle type determines which evaluators can grade it. A reference-answer case can be graded by an exact match; a rubric case cannot. A case whose expected behavior is only described, not turned into a checkable oracle, admits only a human expert. The table below is the full map.
Oracle types
- reference-answer
- One correct output is written down and the answer is compared against it.
- acceptable-set
- Several outputs are acceptable, enumerated or graded, and membership is what is checked.
- rubric
- A written description of what a good answer must do and must not do, adjudicated by a reader or a judge.
- expected-tool-calls
- The actions, not the words: which calls with which arguments, in what order, and which must not happen.
- expected-final-state
- The world after the run, checked by execution: a record, a file, a test that flips from failing to passing.
- expected-behavior-only
- A description of correct behavior that nobody has turned into a checkable oracle yet. Valid to write down, not yet valid to grade against.
Evaluator types
- exact-or-programmatic
- Code decides. String match after a stated normalization, a parser, a schema check, a set membership test.
- rubric-human
- A person scores against a written rubric. Needs at least two raters on a sample to report agreement.
- llm-judge
- A model scores against a rubric, with its own validity evidence against human labels on your items.
- pairwise-preference
- A reader or a judge picks between two outputs. Inadmissible as an oracle-correctness evaluator because it answers which of two outputs is preferred, not which is correct. Where the claim itself is a preference or quality ranking, pairwise is the right instrument.
- trajectory-or-state-check
- The path or the world after: which tools were called, what state changed, which constraint was violated.
- human-expert
- A named domain expert. The only admissible evaluator where correct behavior is a professional judgment.
Which evaluator each oracle type admits
| Oracle type | Admissible evaluators | Admissible with conditions | Inadmissible | Why |
|---|---|---|---|---|
| reference-answer | exact-or-programmatic, rubric-human, human-expert | llm-judge: Admissible only where exact or programmatic comparison cannot express equivalence (free-text answers), and then only as reference-guided grading (the written reference goes in the judge prompt), with agreement evidence against human labels on your own items. Route to judge-validation-report-builder. | pairwise-preference, trajectory-or-state-check | The answer is written down, so a comparison is available that does not need a judgment. A preference between two outputs answers a different question. A trajectory check has no trace to inspect when the oracle is a written answer. |
| acceptable-set | exact-or-programmatic, rubric-human, human-expert | llm-judge: Admissible only where membership requires semantic comparison that exact or programmatic matching cannot express, and then only with agreement evidence against human labels on your own items. Route to judge-validation-report-builder. | pairwise-preference, trajectory-or-state-check | Membership and graded relevance are computable once the set is enumerated. Exact match against one member of the set fails the other members, so the check is membership, never equality. A trajectory check has no trace to inspect. |
| rubric | rubric-human, human-expert | llm-judge: Admissible only with agreement evidence against human labels on your own items. Route to judge-validation-report-builder. | exact-or-programmatic, pairwise-preference, trajectory-or-state-check | Quality here is multi-dimensional and written as criteria. A string comparison cannot read a criterion. A trajectory check inspects the path, not the quality of the output. |
| expected-tool-calls | trajectory-or-state-check, exact-or-programmatic, human-expert | None | llm-judge, pairwise-preference, rubric-human | The calls are recorded, so the check is a comparison against a record. An expert can read a recorded trace. Grading the prose the agent wrote about its actions measures the prose. |
| expected-final-state | trajectory-or-state-check, exact-or-programmatic, human-expert | None | llm-judge, pairwise-preference, rubric-human | The world after the run is checkable by execution. An expert can inspect the resulting state. A judge reading the transcript cannot see the row that was or was not written. |
| expected-behavior-only | human-expert | None | exact-or-programmatic, llm-judge, pairwise-preference, rubric-human, trajectory-or-state-check | Nothing has been turned into a checkable oracle yet. A domain expert can still read the case and say whether the behavior is right; no automated grader can, and running one produces a number about a rule nobody wrote. |
Questions
What does pipe-separated mean in the Slices and Tags columns?
Enter multiple values separated by a vertical bar: refund | happy-path.
The export writes them as an array. Each value is trimmed of surrounding
whitespace, and empty values are dropped.
What happens if I pick an evaluator the oracle type does not admit?
Each oracle type admits a bounded set of evaluators. A reference-answer case can be graded by an exact match, but a judge adds variance the reference removes. A rubric case cannot be graded by a string comparison. The table in "How this works" shows every pairing and the reasoning behind each one. The evaluator dropdown always shows the full list; if you select an inadmissible pairing, the validation panel flags it and the row reads as invalid.
What are the reversibility and blast radius columns?
Reversibility (R0 through R3) describes what happens if the case fails: R0 is no state change, R3 is no way back. Blast radius (B0 through B3) describes how far the effect travels: B0 is one record, B3 is outside the organization. Both are optional and both come from the same scales the action risk matrix uses.
What is provenance?
Where the row came from. Starter rows loaded from a blueprint carry
starter-example. Overlay rows carry
overlay-starter-needs-expert-validation, which means the row
requires domain-expert validation before it can grade anything. Your own rows should carry a provenance that describes their
origin: expert-authored, production-sample,
observed-incident or synthetic-perturbation.
Does refreshing keep my work?
No. The table lives in the page, not in a database. Export before you leave. The canonical JSON and JSONL formats carry every field the table shows. A future version may add local persistence.