LatentEval

INSTRUMENT | Eval statistics

Golden Set Blueprint Builder: Seed a Test Set From a Task Family

Choose a task family, load its case taxonomy and starter rows, then edit, add and export. Every row carries an oracle type and a provenance tag, and the export validates against the eval-set schema.

Start from a task-family blueprint: pick a family, load its case families, slice families and starter rows. Edit them in the table below, add your own rows, and export a dataset that validates against the published eval-set schema with a case profile on every row. The oracle type on each case gates which evaluator can grade it. The structure draws on our research into how agent errors go undetected and what an agent evaluation has to measure. Each task-family pack has its own guide with the reasoning, failure analysis and evaluator guidance behind its rows; this page emits the rows themselves.

Task family

Loading a family replaces the table with its starter rows.

Check this value.

Support correctness is a property of the conversation, not of any single answer. The binding constraint typically appears several turns before the response that must honor it, so a set of isolated single-turn questions evaluates a system nobody actually runs. Three capabilities here are testable and routinely omitted: whether a message satisfying the written complaint definition is recognized as one, whether the system hands off when the situation demands a person, and whether it declines requests it should have handled. The last two move in tandem, so measuring only one tells you nothing about the cost of the other.

Read the Customer support AI blueprint guide (built on how agent errors go undetected, trajectory-level agent evaluation )

Evaluation cases

One row per evaluation case. Oracle type declares what form the expected answer takes. Evaluator says who or what grades the output; the validation panel flags any evaluator the oracle type does not admit. Pipe-separate multiple values in Slices, Tags and Unacceptable. The editor scrolls sideways on a narrow screen.

The table edits the primary fields. Ten secondary fields in the case profile (scenario, actor, evaluationCriteria, linkedConcerns, toolActionExpectation, escalationExpectation, reviewerExpertise, passLogic, version, retirementCondition) are not editable here. To set them, export to JSON or JSONL and edit the file directly.

Case IDInputExpectedSplitSlicesTagsOracle typeExpected behaviorUnacceptableEvaluatorReversibilityBlast radiusProvenanceCase family Actions

Validation

10 rows, 10 valid.

  • Row 1: case.starter-unvalidated - provenance is starter-example, so this case is starting material, not a validated answer.
  • Row 2: case.starter-unvalidated - provenance is starter-example, so this case is starting material, not a validated answer.
  • Row 3: case.starter-unvalidated - provenance is starter-example, so this case is starting material, not a validated answer.
  • Row 4: case.starter-unvalidated - provenance is starter-example, so this case is starting material, not a validated answer.
  • Row 5: case.starter-unvalidated - provenance is starter-example, so this case is starting material, not a validated answer.
  • Row 6: case.starter-unvalidated - provenance is starter-example, so this case is starting material, not a validated answer.
  • Row 7: case.starter-unvalidated - provenance is starter-example, so this case is starting material, not a validated answer.
  • Row 8: case.starter-unvalidated - provenance is starter-example, so this case is starting material, not a validated answer.
  • Row 9: case.starter-unvalidated - provenance is starter-example, so this case is starting material, not a validated answer.
  • Row 10: case.starter-unvalidated - provenance is starter-example, so this case is starting material, not a validated answer.

Export

Each format writes the same rows. Canonical CSV, JSON and JSONL are the eval-set schema itself. The vendor formats reshape the same data for that vendor's harness; the Lossy column below names any field the format cannot carry.

How this works

Every evaluation case declares an oracle type: the form its expected answer takes. The oracle type determines which evaluators can grade it. A reference-answer case can be graded by an exact match; a rubric case cannot. A case whose expected behavior is only described, not turned into a checkable oracle, admits only a human expert. The table below is the full map.

Oracle types

reference-answer
One correct output is written down and the answer is compared against it.
acceptable-set
Several outputs are acceptable, enumerated or graded, and membership is what is checked.
rubric
A written description of what a good answer must do and must not do, adjudicated by a reader or a judge.
expected-tool-calls
The actions, not the words: which calls with which arguments, in what order, and which must not happen.
expected-final-state
The world after the run, checked by execution: a record, a file, a test that flips from failing to passing.
expected-behavior-only
A description of correct behavior that nobody has turned into a checkable oracle yet. Valid to write down, not yet valid to grade against.

Evaluator types

exact-or-programmatic
Code decides. String match after a stated normalization, a parser, a schema check, a set membership test.
rubric-human
A person scores against a written rubric. Needs at least two raters on a sample to report agreement.
llm-judge
A model scores against a rubric, with its own validity evidence against human labels on your items.
pairwise-preference
A reader or a judge picks between two outputs. Inadmissible as an oracle-correctness evaluator because it answers which of two outputs is preferred, not which is correct. Where the claim itself is a preference or quality ranking, pairwise is the right instrument.
trajectory-or-state-check
The path or the world after: which tools were called, what state changed, which constraint was violated.
human-expert
A named domain expert. The only admissible evaluator where correct behavior is a professional judgment.

Which evaluator each oracle type admits

Oracle type Admissible evaluators Admissible with conditions Inadmissible Why
reference-answer exact-or-programmatic, rubric-human, human-expert llm-judge: Admissible only where exact or programmatic comparison cannot express equivalence (free-text answers), and then only as reference-guided grading (the written reference goes in the judge prompt), with agreement evidence against human labels on your own items. Route to judge-validation-report-builder. pairwise-preference, trajectory-or-state-check The answer is written down, so a comparison is available that does not need a judgment. A preference between two outputs answers a different question. A trajectory check has no trace to inspect when the oracle is a written answer.
acceptable-set exact-or-programmatic, rubric-human, human-expert llm-judge: Admissible only where membership requires semantic comparison that exact or programmatic matching cannot express, and then only with agreement evidence against human labels on your own items. Route to judge-validation-report-builder. pairwise-preference, trajectory-or-state-check Membership and graded relevance are computable once the set is enumerated. Exact match against one member of the set fails the other members, so the check is membership, never equality. A trajectory check has no trace to inspect.
rubric rubric-human, human-expert llm-judge: Admissible only with agreement evidence against human labels on your own items. Route to judge-validation-report-builder. exact-or-programmatic, pairwise-preference, trajectory-or-state-check Quality here is multi-dimensional and written as criteria. A string comparison cannot read a criterion. A trajectory check inspects the path, not the quality of the output.
expected-tool-calls trajectory-or-state-check, exact-or-programmatic, human-expert None llm-judge, pairwise-preference, rubric-human The calls are recorded, so the check is a comparison against a record. An expert can read a recorded trace. Grading the prose the agent wrote about its actions measures the prose.
expected-final-state trajectory-or-state-check, exact-or-programmatic, human-expert None llm-judge, pairwise-preference, rubric-human The world after the run is checkable by execution. An expert can inspect the resulting state. A judge reading the transcript cannot see the row that was or was not written.
expected-behavior-only human-expert None exact-or-programmatic, llm-judge, pairwise-preference, rubric-human, trajectory-or-state-check Nothing has been turned into a checkable oracle yet. A domain expert can still read the case and say whether the behavior is right; no automated grader can, and running one produces a number about a rule nobody wrote.

Questions

What does pipe-separated mean in the Slices and Tags columns?

Enter multiple values separated by a vertical bar: refund | happy-path. The export writes them as an array. Each value is trimmed of surrounding whitespace, and empty values are dropped.

What happens if I pick an evaluator the oracle type does not admit?

Each oracle type admits a bounded set of evaluators. A reference-answer case can be graded by an exact match, but a judge adds variance the reference removes. A rubric case cannot be graded by a string comparison. The table in "How this works" shows every pairing and the reasoning behind each one. The evaluator dropdown always shows the full list; if you select an inadmissible pairing, the validation panel flags it and the row reads as invalid.

What are the reversibility and blast radius columns?

Reversibility (R0 through R3) describes what happens if the case fails: R0 is no state change, R3 is no way back. Blast radius (B0 through B3) describes how far the effect travels: B0 is one record, B3 is outside the organization. Both are optional and both come from the same scales the action risk matrix uses.

What is provenance?

Where the row came from. Starter rows loaded from a blueprint carry starter-example. Overlay rows carry overlay-starter-needs-expert-validation, which means the row requires domain-expert validation before it can grade anything. Your own rows should carry a provenance that describes their origin: expert-authored, production-sample, observed-incident or synthetic-perturbation.

Does refreshing keep my work?

No. The table lives in the page, not in a database. Export before you leave. The canonical JSON and JSONL formats carry every field the table shows. A future version may add local persistence.