LatentEval

For builders

Golden set blueprint: structured extraction

Six case families for document extraction, counting omission and hallucination separately and reporting both per-field and per-record metrics.

For builders

In brief

3 POINTS
  • Omission and hallucination fail in different ways, and per-field F1 buries the distinction by collapsing them into one score.
  • Per-record and per-field pass rates diverge: a record with nine correct fields and one fabricated field is 90% right per field and wrong per record.
  • Schema-valid but wrong is the case class you forget, because every one of these failures passes a parse check.

Why this family’s correctness is different

Extraction correctness splits into five classes that per-field F1 collapses: the field is missing (omission), the field is fabricated (hallucination), the span boundary is wrong, the value lands in the wrong field (type-or-label-error), and the output is schema-valid but factually wrong. Per-record and per-field metrics diverge because a record with nine correct fields and one hallucinated one is 90 percent right per field and wrong per record. Report both, with the denominators stated.

This pack builds on research published on this site. How structured output failures happen in practice documents the five error classes and shows why a schema-validation check alone misses most of them. The silent failure problem in AI evaluations explains the general mechanism by which a passing check hides a real error.

Case families

Six families, each testing one of the ways extraction goes wrong. The first two (omission and hallucination) are the ones per-field F1 collapses, and separating them is the point of the split.

These counts seed a starter set; none of them supports a rate claim on its own. The coverage matrix floors set the claimable and readable thresholds.

A set built from this pack carries several oracle types across its families; run the evaluator navigator once per family and write acceptance criteria per family.

TABLEShow full table (6 rows)Showing full table (6 rows)
Case familyWhat it testsExpected behaviorUnacceptable behaviorOracle typeEvaluatorReviewer expertiseSeverity treatmentStarter rows to writeWhat this count licenses
field-omissionWhether the system extracts every required field, not just the easy onesEvery required field is present, with null or a stated reason when the source is silentThe field is silently missing; the field is present with a fabricated valuereference-answerexact-or-programmatic(none)An omitted identifier in a record joined downstream is a different class from an omitted optional note. Count omissions and hallucinations separately15 - 30At 30, worst-case 95% interval: +/-16.8 points
field-hallucinationWhether the system fabricates a value when the source does not carry oneThe field is null or carries a stated reason for absenceA date, name, or number from nowhere in the source; a value inferred from a pattern in other recordsreference-answerexact-or-programmatic(none)A hallucinated identifier or amount inherits the severity of the downstream process that will use it12 - 25At 25, worst-case 95% interval: +/-18.2 points
boundary-errorWhether the extracted span starts or ends in the wrong placeThe span matches the reference exactly, by character offset or token boundaryThe span includes extra context; the span cuts a word in the middlereference-answerexact-or-programmatic(none)Usually low severity unless the extra or missing characters change the meaning10 - 20At 20, worst-case 95% interval: +/-20.1 points
type-or-label-errorWhether the value lands in the correct schema fieldEvery value is assigned to the field the schema declares for itA date in the amount field; a shipping address in the billing fieldreference-answerexact-or-programmatic(none)A wrong assignment the downstream consumer treats as correct data is high severity. A wrong type the parser rejects is loud and cheap10 - 20At 20, worst-case 95% interval: +/-20.1 points
schema-valid-wrongWhether the value is what the document says, not just whether it fits the schemaThe field value matches the reference from the source documentA value from a different section of the same document; a value that is schema-valid and factually wrongreference-answerexact-or-programmatic(none)This is the class people forget, because a parse check passes every one of them. Severity follows the field and the downstream process12 - 25At 25, worst-case 95% interval: +/-18.2 points
multi-record-extractionWhether the system extracts all records from a multi-record document without merging or droppingOne output record per source record, no records merged, no records droppedTwo records merged into one; a record dropped because it looked like a duplicateexpected-final-stateexact-or-programmatic(none)A dropped record is counted against the per-record denominator. A merged record is its own error class8 - 15At 15, worst-case 95% interval: +/-22.6 points

schema-valid-wrong is the case class people forget. A pipeline that validates output against the schema and reports green has told you nothing about whether the extracted values are correct. The parse check passes every one of these failures.

Slice families

TABLEShow full table (5 rows)Showing full table (5 rows)
Slice familyWhy it mattersExample valuesSensitive
Content or document typeAn invoice, a contract, and a clinical note carry different structures and different failure shapesinvoice, contract, clinical note, receipt, formNo
Formatting shiftA table, a free-text paragraph, and a scanned image of the same data fail in different waysstructured table, free-text paragraph, OCR output, PDF extractNo
Input lengthExtraction from a one-page invoice and from a 40-page contract fail at different ratesunder 500 words, 500 to 5,000 words, over 5,000 wordsNo
Upstream sourceAn OCR scan, a native PDF, and a typed form carry different noise levelsnative PDF, OCR scan, typed form, email attachmentNo
Language and localeDate formats, number separators, and field labels differ by locale, each a source of extraction errorsen-US, de-DE, ja-JP, mixed-language documentNo

None of these slices is marked sensitive. The formatting-shift slice is particularly important: the same document in a structured table and in free-text paragraph form fails at different rates and in different ways, and a set built on only one layout says nothing about the other.

Coverage matrix skeleton

Each cell names what you would not know if the cell were empty. Use the eval coverage matrix tool to fill it against your own system.

TABLEShow full table (11 rows)Showing full table (11 rows)
Case familySlice familyWhat an empty cell hides
field-omissionContent or document typeWhether omission rates vary by document type
field-omission(overall)Whether the system silently drops required fields
field-hallucinationContent or document typeWhether the system fabricates values differently by document type
field-hallucinationUpstream sourceWhether hallucination rates are higher on OCR scans
boundary-errorFormatting shiftWhether span boundaries break more often on certain layouts
type-or-label-error(overall)Whether the system assigns values to the wrong fields
type-or-label-errorLanguage and localeWhether label assignment breaks on non-English documents
schema-valid-wrongContent or document typeWhether the system passes schema checks and is factually wrong
schema-valid-wrongInput lengthWhether schema-valid errors increase with document length
multi-record-extractionFormatting shiftWhether the system merges or drops records in certain layouts
multi-record-extraction(overall)Whether the system extracts every record from a multi-record document

Evaluator and annotation guidance

All six case families use programmatic comparison. The split is between reference-answer (five families) and expected-final-state (multi-record-extraction).

Reference-answer cases (exact-or-programmatic)

Used for field-omission, field-hallucination, boundary-error, type-or-label-error, and schema-valid-wrong.

Validity evidence needed:

  • The normalization is written down before anything is graded: whitespace, case, number format, date format, and what counts as a match for each field type.
  • Per-field and per-record metrics are both computed, and the denominators are stated.
  • A sample of the passes and the failures is read by a person once.

Per-field and per-record metrics diverge. A record with nine correct fields and one hallucinated one is 90 percent right per field and wrong per record. Reporting only the per-field number tells you nothing about how many records a downstream process will act on incorrectly. Report both.

Metrics: pass rate with a confidence interval, per-slice pass rate.

Tools: pass rate CI calculator, slice and class balance analyzer.

Expected-final-state cases (exact-or-programmatic)

Used for multi-record-extraction, where the oracle is the set of records the document contains.

Validity evidence needed:

  • The record count is checked before the field values, so a merged record fails immediately.
  • The fixture resets between runs.

Metrics: pass rate, constraint violation count.

Tools: pass rate CI calculator.

Starter rows

These rows are starter examples showing what a case in this family looks like. They are not validated gold-standard answers. Before any of them grades a real output, someone who knows your system needs to confirm or rewrite the expected behavior in each one.

TABLEShow full table (10 rows)Showing full table (10 rows)
IDInput (abbreviated)Case familyOracle typeEvaluatorReversibilityBlast radius
ext-001Extract invoice number, date, total, vendor from an invoicefield-omission / hallucination (baseline)reference-answerexact-or-programmaticR0B0
ext-002Extract vendor name and PO number; vendor name is absentfield-hallucinationreference-answerexact-or-programmaticR1B0
ext-003Extract patient name and DOB; two names appear (patient and physician)type-or-label-errorreference-answerexact-or-programmaticR2B2
ext-004Extract three line items; two share a name but are distinctmulti-record-extractionexpected-final-stateexact-or-programmaticR1B1
ext-005Extract start date and end date from a contract clausetype-or-label-errorreference-answerexact-or-programmaticR1B1
ext-006Extract the total from a document with subtotal, discount, tax, and totalschema-valid-wrongreference-answerexact-or-programmaticR1B1
ext-007Extract clause number 4.2(a)(iii) with nested parentheticalsboundary-errorreference-answerexact-or-programmaticR0B0
ext-008Extract shipping address when billing and shipping appear side by sidetype-or-label-errorreference-answerexact-or-programmaticR2B1
ext-009Extract medication name and dosage from a prescription linefield-omissionreference-answerexact-or-programmaticR2B2
ext-010Extract company name from a title block where the name is at the bottomschema-valid-wrongreference-answerexact-or-programmaticR0B0

Failure-to-case mapping

Case familyConcern presets caught
field-omissionomission, format-contract-violation
field-hallucinationunsupported-claim, format-contract-violation
boundary-errorformat-contract-violation, omission
type-or-label-errorwrong-entity, format-contract-violation
schema-valid-wrongunsupported-claim, wrong-entity
multi-record-extractionomission, wrong-entity

Use the risk-to-test mapper to trace from a named concern to the case families that test for it.

Downloads

Starter files for this pack, generated from the same source the examples above are read from: the CSV and the JSONL. Both carry the case profile in metadata.case. Every row is starting material, not a validated answer. This family’s files are published as structured-extraction.csv and structured-extraction.jsonl. To edit them in the browser and add your own, use the blueprint builder.

Which instruments apply

Companion pack guides

  • Customer support covers multi-turn agent conversations with case families for policy lookup, escalation, and complaint recognition.
  • Transactional agents tests agents that execute real-world write actions, where extraction accuracy feeds directly into whether the correct action fires.