For builders
Golden set blueprint: structured extraction
Six case families for document extraction, counting omission and hallucination separately and reporting both per-field and per-record metrics.
In brief
3 POINTS- Omission and hallucination fail in different ways, and per-field F1 buries the distinction by collapsing them into one score.
- Per-record and per-field pass rates diverge: a record with nine correct fields and one fabricated field is 90% right per field and wrong per record.
- Schema-valid but wrong is the case class you forget, because every one of these failures passes a parse check.
On this page (10)
Why this family’s correctness is different
Extraction correctness splits into five classes that per-field F1 collapses: the field is missing (omission), the field is fabricated (hallucination), the span boundary is wrong, the value lands in the wrong field (type-or-label-error), and the output is schema-valid but factually wrong. Per-record and per-field metrics diverge because a record with nine correct fields and one hallucinated one is 90 percent right per field and wrong per record. Report both, with the denominators stated.
This pack builds on research published on this site. How structured output failures happen in practice documents the five error classes and shows why a schema-validation check alone misses most of them. The silent failure problem in AI evaluations explains the general mechanism by which a passing check hides a real error.
Case families
Six families, each testing one of the ways extraction goes wrong. The first two (omission and hallucination) are the ones per-field F1 collapses, and separating them is the point of the split.
These counts seed a starter set; none of them supports a rate claim on its own. The coverage matrix floors set the claimable and readable thresholds.
A set built from this pack carries several oracle types across its families; run the evaluator navigator once per family and write acceptance criteria per family.
TABLEShow full table (6 rows)Showing full table (6 rows)
| Case family | What it tests | Expected behavior | Unacceptable behavior | Oracle type | Evaluator | Reviewer expertise | Severity treatment | Starter rows to write | What this count licenses |
|---|---|---|---|---|---|---|---|---|---|
| field-omission | Whether the system extracts every required field, not just the easy ones | Every required field is present, with null or a stated reason when the source is silent | The field is silently missing; the field is present with a fabricated value | reference-answer | exact-or-programmatic | (none) | An omitted identifier in a record joined downstream is a different class from an omitted optional note. Count omissions and hallucinations separately | 15 - 30 | At 30, worst-case 95% interval: +/-16.8 points |
| field-hallucination | Whether the system fabricates a value when the source does not carry one | The field is null or carries a stated reason for absence | A date, name, or number from nowhere in the source; a value inferred from a pattern in other records | reference-answer | exact-or-programmatic | (none) | A hallucinated identifier or amount inherits the severity of the downstream process that will use it | 12 - 25 | At 25, worst-case 95% interval: +/-18.2 points |
| boundary-error | Whether the extracted span starts or ends in the wrong place | The span matches the reference exactly, by character offset or token boundary | The span includes extra context; the span cuts a word in the middle | reference-answer | exact-or-programmatic | (none) | Usually low severity unless the extra or missing characters change the meaning | 10 - 20 | At 20, worst-case 95% interval: +/-20.1 points |
| type-or-label-error | Whether the value lands in the correct schema field | Every value is assigned to the field the schema declares for it | A date in the amount field; a shipping address in the billing field | reference-answer | exact-or-programmatic | (none) | A wrong assignment the downstream consumer treats as correct data is high severity. A wrong type the parser rejects is loud and cheap | 10 - 20 | At 20, worst-case 95% interval: +/-20.1 points |
| schema-valid-wrong | Whether the value is what the document says, not just whether it fits the schema | The field value matches the reference from the source document | A value from a different section of the same document; a value that is schema-valid and factually wrong | reference-answer | exact-or-programmatic | (none) | This is the class people forget, because a parse check passes every one of them. Severity follows the field and the downstream process | 12 - 25 | At 25, worst-case 95% interval: +/-18.2 points |
| multi-record-extraction | Whether the system extracts all records from a multi-record document without merging or dropping | One output record per source record, no records merged, no records dropped | Two records merged into one; a record dropped because it looked like a duplicate | expected-final-state | exact-or-programmatic | (none) | A dropped record is counted against the per-record denominator. A merged record is its own error class | 8 - 15 | At 15, worst-case 95% interval: +/-22.6 points |
schema-valid-wrong is the case class people forget. A pipeline that validates output against the schema and reports green has told you nothing about whether the extracted values are correct. The parse check passes every one of these failures.
Slice families
TABLEShow full table (5 rows)Showing full table (5 rows)
| Slice family | Why it matters | Example values | Sensitive |
|---|---|---|---|
| Content or document type | An invoice, a contract, and a clinical note carry different structures and different failure shapes | invoice, contract, clinical note, receipt, form | No |
| Formatting shift | A table, a free-text paragraph, and a scanned image of the same data fail in different ways | structured table, free-text paragraph, OCR output, PDF extract | No |
| Input length | Extraction from a one-page invoice and from a 40-page contract fail at different rates | under 500 words, 500 to 5,000 words, over 5,000 words | No |
| Upstream source | An OCR scan, a native PDF, and a typed form carry different noise levels | native PDF, OCR scan, typed form, email attachment | No |
| Language and locale | Date formats, number separators, and field labels differ by locale, each a source of extraction errors | en-US, de-DE, ja-JP, mixed-language document | No |
None of these slices is marked sensitive. The formatting-shift slice is particularly important: the same document in a structured table and in free-text paragraph form fails at different rates and in different ways, and a set built on only one layout says nothing about the other.
Coverage matrix skeleton
Each cell names what you would not know if the cell were empty. Use the eval coverage matrix tool to fill it against your own system.
TABLEShow full table (11 rows)Showing full table (11 rows)
| Case family | Slice family | What an empty cell hides |
|---|---|---|
| field-omission | Content or document type | Whether omission rates vary by document type |
| field-omission | (overall) | Whether the system silently drops required fields |
| field-hallucination | Content or document type | Whether the system fabricates values differently by document type |
| field-hallucination | Upstream source | Whether hallucination rates are higher on OCR scans |
| boundary-error | Formatting shift | Whether span boundaries break more often on certain layouts |
| type-or-label-error | (overall) | Whether the system assigns values to the wrong fields |
| type-or-label-error | Language and locale | Whether label assignment breaks on non-English documents |
| schema-valid-wrong | Content or document type | Whether the system passes schema checks and is factually wrong |
| schema-valid-wrong | Input length | Whether schema-valid errors increase with document length |
| multi-record-extraction | Formatting shift | Whether the system merges or drops records in certain layouts |
| multi-record-extraction | (overall) | Whether the system extracts every record from a multi-record document |
Evaluator and annotation guidance
All six case families use programmatic comparison. The split is between reference-answer (five families) and expected-final-state (multi-record-extraction).
Reference-answer cases (exact-or-programmatic)
Used for field-omission, field-hallucination, boundary-error, type-or-label-error, and schema-valid-wrong.
Validity evidence needed:
- The normalization is written down before anything is graded: whitespace, case, number format, date format, and what counts as a match for each field type.
- Per-field and per-record metrics are both computed, and the denominators are stated.
- A sample of the passes and the failures is read by a person once.
Per-field and per-record metrics diverge. A record with nine correct fields and one hallucinated one is 90 percent right per field and wrong per record. Reporting only the per-field number tells you nothing about how many records a downstream process will act on incorrectly. Report both.
Metrics: pass rate with a confidence interval, per-slice pass rate.
Tools: pass rate CI calculator, slice and class balance analyzer.
Expected-final-state cases (exact-or-programmatic)
Used for multi-record-extraction, where the oracle is the set of records the document contains.
Validity evidence needed:
- The record count is checked before the field values, so a merged record fails immediately.
- The fixture resets between runs.
Metrics: pass rate, constraint violation count.
Tools: pass rate CI calculator.
Starter rows
These rows are starter examples showing what a case in this family looks like. They are not validated gold-standard answers. Before any of them grades a real output, someone who knows your system needs to confirm or rewrite the expected behavior in each one.
TABLEShow full table (10 rows)Showing full table (10 rows)
| ID | Input (abbreviated) | Case family | Oracle type | Evaluator | Reversibility | Blast radius |
|---|---|---|---|---|---|---|
| ext-001 | Extract invoice number, date, total, vendor from an invoice | field-omission / hallucination (baseline) | reference-answer | exact-or-programmatic | R0 | B0 |
| ext-002 | Extract vendor name and PO number; vendor name is absent | field-hallucination | reference-answer | exact-or-programmatic | R1 | B0 |
| ext-003 | Extract patient name and DOB; two names appear (patient and physician) | type-or-label-error | reference-answer | exact-or-programmatic | R2 | B2 |
| ext-004 | Extract three line items; two share a name but are distinct | multi-record-extraction | expected-final-state | exact-or-programmatic | R1 | B1 |
| ext-005 | Extract start date and end date from a contract clause | type-or-label-error | reference-answer | exact-or-programmatic | R1 | B1 |
| ext-006 | Extract the total from a document with subtotal, discount, tax, and total | schema-valid-wrong | reference-answer | exact-or-programmatic | R1 | B1 |
| ext-007 | Extract clause number 4.2(a)(iii) with nested parentheticals | boundary-error | reference-answer | exact-or-programmatic | R0 | B0 |
| ext-008 | Extract shipping address when billing and shipping appear side by side | type-or-label-error | reference-answer | exact-or-programmatic | R2 | B1 |
| ext-009 | Extract medication name and dosage from a prescription line | field-omission | reference-answer | exact-or-programmatic | R2 | B2 |
| ext-010 | Extract company name from a title block where the name is at the bottom | schema-valid-wrong | reference-answer | exact-or-programmatic | R0 | B0 |
Failure-to-case mapping
| Case family | Concern presets caught |
|---|---|
| field-omission | omission, format-contract-violation |
| field-hallucination | unsupported-claim, format-contract-violation |
| boundary-error | format-contract-violation, omission |
| type-or-label-error | wrong-entity, format-contract-violation |
| schema-valid-wrong | unsupported-claim, wrong-entity |
| multi-record-extraction | omission, wrong-entity |
Use the risk-to-test mapper to trace from a named concern to the case families that test for it.
Downloads
Starter files for this pack, generated from the same source the examples above are read from: the CSV and the JSONL. Both carry the case profile in metadata.case. Every row is starting material, not a validated answer. This family’s files are published as structured-extraction.csv and structured-extraction.jsonl. To edit them in the browser and add your own, use the blueprint builder.
Which instruments apply
- Pass rate CI calculator for a confidence interval on the pass rate per family and per record versus per field.
- Slice and class balance analyzer for checking whether the set covers each case-by-slice cell.
- Eval coverage matrix for the grid of which cells are populated and which are empty.
- Golden set blueprint builder for editing and extending the starter rows.
- Eval dataset schema validator for checking that exported rows validate against the eval-set schema.
- Risk-to-test mapper for mapping from a named concern to the case families that test for it.
Companion pack guides
- Customer support covers multi-turn agent conversations with case families for policy lookup, escalation, and complaint recognition.
- Transactional agents tests agents that execute real-world write actions, where extraction accuracy feeds directly into whether the correct action fires.