For builders
Evaluation case schema, field by field
The case profile published atop the eval-set schema: what each field does, why the oracle type gates the evaluator, what is left out on purpose, and how rows round-trip.
In brief
4 POINTS- Each case declares its oracle type, and the oracle type determines which evaluators are admissible.
- Most golden sets fail because oracle types get mixed in one column and a single grader gets applied to all of them.
- Four fields are required, not fifteen: a wide flat schema stays mostly empty and invites the mixing it exists to prevent.
- There is no difficulty, no weight, no priority, and no per-case score: a weight column is a composite waiting for arithmetic nobody can justify.
On this page (8)
Most golden sets that fail in practice fail for the same reason: someone writes a spreadsheet where every row has an input, an expected output, and a column that says “score.” The grader reads the column, compares the output, and returns a number. The number is wrong, and nobody finds out until someone asks why a model that scores 92% just refunded a customer who never bought anything.
The problem is not the spreadsheet. The problem is that the column conflates several different kinds of expected answer, and the grader treats them identically. A row whose expected answer is a verbatim string and a row whose expected behavior is a written description of what a competent agent would do are not the same kind of case. They require different evaluators, and the evaluator that works for one produces a meaningless number for the other.
The case profile published in the eval-set schema is a convention that makes these distinctions explicit. It lives at metadata.case on every row, and it carries a published JSON Schema that the validator checks against. What follows is the schema field by field, starting with the design rule at its center.
The oracle type gates the evaluator
Every evaluation case declares an oracle type: the form its expected answer takes. Six are recognized.
reference-answer. One correct output is written down and the answer is compared against it. This is the strongest oracle: it removes the need for judgment.
acceptable-set. Several outputs are acceptable, enumerated or graded, and the check is membership, not equality. An exact match against one member of the set fails the other members, so the evaluator must check set membership.
rubric. A written description of what a good answer must do and must not do, adjudicated by a reader or a model judge. Quality here is multi-dimensional, and a string comparison cannot read a criterion.
expected-tool-calls. The actions, not the words: which calls with which arguments, in what order, and which must not happen. The evaluator reads the recorded trajectory, never the prose the agent wrote about its actions.
expected-final-state. The world after the run, checked by execution: a record, a file, a test that flips from failing to passing. A judge reading the transcript cannot see the row that was or was not written.
expected-behavior-only. A description of correct behavior that nobody has turned into a checkable oracle yet. This is a legitimate thing to write down, but it is not yet a legitimate thing to grade against. The only admissible evaluator is a human expert who can read the case and say whether the behavior is right; no automated grader can, and running one produces a number about a rule nobody wrote.
Each oracle type admits a bounded set of evaluators and bars the rest. The blueprint builder enforces these pairings, and the validator flags any row that names an inadmissible evaluator for its oracle type. The full map is in the next section.
Admissible evaluators
Six evaluator types are recognized: exact-or-programmatic (code decides), rubric-human (a person scores against a rubric), llm-judge (a model scores against a rubric, with validity evidence), pairwise-preference (a reader or judge picks between two outputs), trajectory-or-state-check (the path or the world after), and human-expert (a named domain expert).
The admissibility map:
- reference-answer admits exact-or-programmatic, rubric-human, and human-expert. Admits llm-judge with conditions: the reference goes in the judge prompt (reference-guided grading), and the judge is validated against human labels on a sample of the task’s own cases (route to judge-validation-report-builder). Bars pairwise-preference (a preference between two outputs answers a different question) and trajectory-or-state-check (the oracle is a written answer, not a trace).
- acceptable-set admits exact-or-programmatic, rubric-human, and human-expert. Admits llm-judge with conditions: the acceptable set goes in the judge prompt (reference-guided grading), and the judge is validated against human labels on a sample of the task’s own cases (route to judge-validation-report-builder). Bars pairwise-preference (the check is membership, never which of two outputs a reader prefers) and trajectory-or-state-check (there is no trace to inspect).
- rubric admits rubric-human and human-expert. Admits llm-judge with conditions: the rubric goes in the judge prompt, and the judge is validated against human labels on a sample of the task’s own cases (route to judge-validation-report-builder). Bars exact-or-programmatic (a string comparison cannot read a criterion), pairwise-preference (a preference is not a rubric score), and trajectory-or-state-check (the evaluator inspects the path, not the quality of the output).
- expected-tool-calls admits trajectory-or-state-check, exact-or-programmatic, and human-expert. Bars llm-judge, pairwise-preference, and rubric-human: the calls are recorded, so the check is a comparison against a record. An expert can read a recorded trace.
- expected-final-state admits trajectory-or-state-check, exact-or-programmatic, and human-expert. Bars llm-judge, pairwise-preference, and rubric-human: the world after the run is checkable by execution. An expert can inspect the resulting state.
- expected-behavior-only admits human-expert only. Bars everything else: nothing has been turned into a checkable oracle yet.
This is not an imported taxonomy. It is a synthesis drawn from several public evaluation formats and stated by none of them. The reasoning behind each pairing is printed in the blueprint builder and in the validator.
Required fields
Four fields are required; every other field in the schema is optional. A wide flat schema that asks for fifteen columns on every row is mostly empty and invites the mixing it was meant to prevent.
caseId (string). A unique identifier for this case. The validator checks that it matches the row’s top-level id field; a mismatch is flagged. Use a slug that makes the case recognizable in a list: cs-refund-valid-receipt, not row-042.
oracleType (enum). One of the six oracle types listed above. This is the field that gates which evaluator the row admits. Setting it is the single most consequential decision in writing the case.
expectedBehavior (string). What the agent should do, in the author’s own words. This is not the expected output; it is a description of the behavior the expected output is meant to represent. A reference-answer case still carries a behavior description alongside the literal answer: the description says why the answer is right.
provenance (enum). Where the row came from: observed-incident, production-sample, expert-authored, starter-example, overlay-starter-needs-expert-validation, synthetic-perturbation, or imported. Every row carries its origin, and the validator treats the two starter provenances differently: they mark published starting material, and every surface that renders such a row says that it is an example, not a validated gold-standard answer.
Optional fields
These fields let a case carry more context when the author has it. None is required, and none changes which evaluator the row admits.
scenario (string). The situation the case describes, in enough detail that a reader who has never seen the system under test can understand what the case is testing.
actor (string). Who or what is performing the action in the case: a customer, an admin, an internal API.
unacceptableBehavior (string array). Behaviors the output must not exhibit. Pipe-separated in CSV. A case with both an expected behavior and an unacceptable behavior is making two claims: one about what should happen and one about what must not.
evaluationCriteria (string array). The specific criteria against which the output is evaluated, if the oracle type is rubric. A rubric row without evaluation criteria is incomplete but not invalid; the criteria can live in the rubric itself rather than in this field.
reversibility (enum: R0 through R3). What happens if the case fails. R0 is no state change. R1 is the actor can undo it. R2 requires a compensating operation. R3 is no way back. These are the same levels the action risk matrix uses.
blastRadius (enum: B0 through B3). How far the effect of a failure travels. B0 is one record. B1 is one team. B2 is the organization. B3 is outside. Same scale as the risk matrix’s blast radius.
linkedConcerns (string array). Named evaluation concerns this case covers, linking it to the failure concerns the risk-to-test mapper works from. Those presets are a starting palette, not a published taxonomy.
evaluator (evaluator type enum). Which evaluator will grade this case. If named, the validator checks it against the oracle type and flags inadmissible pairings. If absent, any admissible evaluator for the oracle type is implied.
toolActionExpectation (string). For tool-calling cases: which tools should be called, with which arguments.
escalationExpectation (string). Whether and when the case should escalate to a human, and what the escalation should look like.
reviewerExpertise (string). What domain expertise the reviewer needs to judge this case correctly. A case graded by a human-expert evaluator should always state this.
passLogic (string). How the evaluator’s output maps to a pass or fail: the threshold, the logic, the conditions. A rubric case might say “3 of 5 criteria met”; a reference-answer case might say “exact match after lowercasing.”
version (string). The case’s own version, separate from the set version. Changed when the expected behavior changes.
retirementCondition (string). When this case should be removed from the set: after a model reaches a threshold, after a feature ships, after a date.
caseFamily (string). The case-family id this row belongs to within its task family. A case family groups rows that test the same kind of failure: “policy-lookup” in a customer-support set, “field-omission” in an extraction set. The id is the kebab-case key from the pack’s case-families array. Rows that share a case family share the same oracle type, the same evaluator, and the same severity treatment; the blueprint builder populates this field automatically when loading starter rows.
extensions (object). A namespace for fields the schema does not cover. Each key is a namespace string; each value is an object. Extensions are carried through export and import but not validated beyond their shape.
What the schema deliberately omits
No difficulty. A difficulty column is a weight by another name, and a weight column is a composite waiting for an arithmetic. A set with difficulty 1 through 5 and a score formula that weights them has quietly built a composite index, and the composite hides every design decision the case profile is meant to surface.
No weight. Same reasoning. If some cases matter more than others, that is a property of the eval harness’s scoring function, not of the case.
No priority. Priority is a triage concept, not a property of a test case. A P1 column tempts the set author into marking half the cases P1, which tells the harness nothing.
No per-case score. The case profile describes what the case tests and how to grade it. The score is the grader’s output, not the case’s input.
How it round-trips
The case profile lives at metadata.case on every row. In the canonical CSV, case-profile fields flatten to columns prefixed case.: case.oracleType, case.expectedBehavior, case.reversibility. List fields (unacceptableBehavior, evaluationCriteria, linkedConcerns) pipe-separate their values with a | delimiter. The canonical JSON and JSONL carry the full nested structure under the metadata object.
The eval-dataset schema validator checks every row’s case profile against the published JSON Schema, checks the oracle-evaluator admissibility rule, and checks that caseId matches the row’s top-level id. The blueprint builder exports in every canonical format plus vendor formats for Promptfoo, DeepEval, LangSmith, and Braintrust; the validator accepts all of them on import.
Every vendor format loses some fields because the vendor’s schema does not have a place for them. The lossy property on an export result tells you which fields were dropped and why. The canonical formats are lossless: a row that round-trips through canonical CSV, JSON, or JSONL comes back identical.
The mixing problem
The most common failure in golden-set construction is not missing cases. It is mixing oracle types without noticing. A spreadsheet with a column called “expected” that holds a verbatim string in row 3 and a behavioral description in row 7 has silently asked the grader to do two different things in the same column. An exact-match grader applied to both rows will correctly grade row 3 and produce a meaningless result for row 7. The meaningless result will be a number, and the number will be included in the aggregate.
The case profile makes this visible. When every row declares its oracle type, a grader configuration that applies exact-or-programmatic to a rubric case is a type error, caught by the validator before any model runs. The cost of the declaration is one field per row. The cost of not declaring it is a pass rate that includes rows where the grader answered a question nobody asked.
This is also why the evaluator field is optional rather than required. An author who sets oracleType correctly has already constrained the evaluator to the admissible set. Naming the evaluator explicitly is useful for documentation and for harness configuration, but the admissibility check works whether the evaluator is named or not: if it is named and inadmissible, the validator catches it. If it is absent, the harness is free to use any admissible evaluator.
Building a set from this schema
Start with the blueprint builder. Pick a task family, load its starter rows, and edit them. The starter rows carry provenance: starter-example, which means they are shaped correctly but their expected behaviors are examples, not validated gold-standard answers. Your first job is to replace those expected behaviors with answers that are correct for your system.
Then set the oracle type on every row, and let the tool narrow the evaluator. If you have cases whose expected answer is a verbatim string, those are reference-answer cases. If you have cases whose expected behavior is a multi-dimensional rubric, those are rubric cases. If you have cases where the right answer is “the agent should call this tool with these arguments,” those are expected-tool-calls cases. The oracle type is not a property of your system; it is a property of how you plan to check whether the system did the right thing.
Add your own cases. Set provenance to match how you got them: production-sample if you pulled them from logs, expert-authored if a domain expert wrote them, observed-incident if they came from a failure that actually happened. The provenance field does not change how the case is graded, but it tells anyone reading the set how the expected behavior was established, which is what separates a golden set from a collection of inputs with guesses.