For builders
Golden set blueprint: summarization
Six case families for summarization: source contradictions, unsupported additions, and salient omissions counted separately, with evaluator granularity tested as a variable.
In brief
3 POINTS- A contradiction and an unsupported addition are different failures, and the distinction holds only once you name a source for each claim.
- Evaluator granularity is itself a variable in the result: a document-level check misses errors that a sentence-level check catches.
- You score omissions against what the summary was asked to preserve, not against the full length of the source.
On this page (10)
Why this family’s correctness is different
A faithfulness number that does not say whether the error was a contradiction, an unsupported addition, or a dropped finding is a number that cannot point you at the fix. Intrinsic errors contradict the source; extrinsic errors add something the source never said; omissions drop something the source emphasized. The distinction only means anything once a source is named for each claim. And the evaluator granularity matters: a document-level consistency check misses errors a sentence-level check catches, so the granularity of the evaluator is itself a variable in the result.
This pack builds on research published on this site. The silent failure problem in AI evaluations documents how a single faithfulness number hides the failure shape. How LLM-as-a-judge evaluations work and where they break covers the granularity problem: a judge applied at the wrong level of analysis adds noise that looks like signal. A practical framework for AI agent evaluation establishes the general principle that each failure type needs its own denominator.
Case families
Six families. The first three (intrinsic-contradiction, extrinsic-addition, salient-omission) separate what a single faithfulness number collapses. The fourth (granularity-mismatch) catches the case where the summary does not match the requested detail level. The fifth (entity-confusion) catches the case where the summary attributes a fact to the wrong entity. The sixth (evaluator-granularity-failure) is a meta-case: it tests whether the chosen evaluator operates at a fine enough granularity to catch the errors the other families describe.
These counts seed a starter set; none of them supports a rate claim on its own. The coverage matrix floors set the claimable and readable thresholds.
A set built from this pack carries several oracle types across its families; run the evaluator navigator once per family and write acceptance criteria per family.
TABLEShow full table (6 rows)Showing full table (6 rows)
| Case family | What it tests | Expected behavior | Unacceptable behavior | Oracle type | Evaluator | Reviewer expertise | Severity treatment | Starter rows to write | What this count licenses |
|---|---|---|---|---|---|---|---|---|---|
| intrinsic-contradiction | Whether the summary states something the source explicitly says differently | Every claim is consistent with the corresponding passage in the source | States a number, date, or name that differs; reverses a causal direction the source states | rubric | rubric-human | Someone who can spot a factual conflict between the summary and the source | Severity follows the claim. Count contradictions separately from extrinsic errors | 15 - 30 | At 30, worst-case 95% interval: +/-16.8 points |
| extrinsic-addition | Whether the summary introduces a claim the source is silent on, true or not | Every claim traces to a passage in the source or is explicitly marked as external | States a fact the source does not contain without flagging it; infers a conclusion and presents it as the source’s finding | rubric | rubric-human | Someone who can confirm whether a claim is supported, not just plausible | An extrinsic-addition that happens to be true in the world is still an error against the source | 12 - 25 | At 25, worst-case 95% interval: +/-18.2 points |
| salient-omission | Whether the summary preserves the points the source emphasizes | Every conclusion, finding, warning, or recommendation the source calls out appears | Drops a safety warning; covers positive findings and omits the limitations | rubric | rubric-human | Someone who can decide which points are salient, not just present | Count omissions against a stated preservation target, not an abstract completeness goal | 12 - 25 | At 25, worst-case 95% interval: +/-18.2 points |
| granularity-mismatch | Whether the summary matches the level of detail the request asks for | Detail level matches: a one-paragraph summary does not itemize every finding | A detailed summary when one paragraph was asked for; three sentences when comprehensive was asked for | rubric | rubric-human | Someone who can judge whether the detail level serves the stated purpose | Usually low severity on its own. Rises when the wrong granularity causes an omission | 8 - 15 | At 15, worst-case 95% interval: +/-22.6 points |
| entity-confusion | Whether the summary attributes a fact to the right entity when the source discusses several | Every attribution matches the entity the source assigns it to | Attributes company A’s figure to company B; merges two people’s statements into one | reference-answer | exact-or-programmatic | (none) | A misattributed figure carries the severity of the decision it feeds | 8 - 15 | At 15, worst-case 95% interval: +/-22.6 points |
| evaluator-granularity-failure | Whether the chosen evaluator operates at a granularity fine enough to detect the error class | The evaluation catches the seeded error at the sentence or claim level | A document-level check passes a summary that contradicts one sentence; NLI at full-document level misses sentence-level inconsistency | rubric | llm-judge | Someone who can compare sentence-level and document-level verdicts | A passed check that should have failed undermines the numbers for every other case | 6 - 12 | At 12, worst-case 95% interval: +/-24.6 points |
The evaluator-granularity-failure family is a meta-case. It does not test the summarizer; it tests the evaluator. A document-level faithfulness check that passes a summary containing a sentence-level contradiction has not measured faithfulness at the resolution that matters. Including these cases in the set means the set can detect when its own evaluation is too coarse.
Slice families
TABLEShow full table (5 rows)Showing full table (5 rows)
| Slice family | Why it matters | Example values | Sensitive |
|---|---|---|---|
| Content or document type | A clinical abstract, a legal contract, and a meeting transcript carry different structures and error shapes | clinical abstract, legal contract, earnings report, meeting transcript, policy document | No |
| Input length | Omissions rise with source length. A set built on short documents says nothing about a system that summarizes long ones | under 1,000 words, 1,000 to 5,000 words, over 5,000 words | No |
| Intent or task type | An executive summary, a comparative summary, and a safety-critical summary have different preservation targets and failure costs | executive summary, detailed summary, comparative, section-scoped, action-item extraction | No |
| Language and locale | Summarization quality often differs by language, and a cross-language summary adds a translation failure surface | en-US, de-DE, ja-JP, cross-language | No |
| Upstream source | An OCR-scanned PDF and a clean digital document carry different noise levels, and extraction errors in the source propagate into the summary | clean digital document, OCR scan, speech-to-text transcript, web scrape | No |
The input-length slice deserves special attention. Omissions are the error class most sensitive to source length: a system that omits nothing from a 500-word abstract may omit material findings from a 5,000-word report. A set built entirely on short sources has a structural blind spot.
Coverage matrix skeleton
Use the eval coverage matrix tool to fill the grid against your own system.
TABLEShow full table (11 rows)Showing full table (11 rows)
| Case family | Slice family | What an empty cell hides |
|---|---|---|
| intrinsic-contradiction | Content or document type | Whether contradiction rates differ by document type |
| intrinsic-contradiction | (overall) | Whether the summarizer contradicts the source at all |
| extrinsic-addition | Content or document type | Whether the summarizer fabricates more on one document type than another |
| extrinsic-addition | (overall) | Whether the summarizer adds unsupported claims |
| salient-omission | Input length | Whether omissions rise with source length |
| salient-omission | (overall) | Whether the summarizer drops important points |
| granularity-mismatch | Intent or task type | Whether the system matches the requested detail level |
| entity-confusion | (overall) | Whether the summarizer swaps entities when the source discusses several |
| evaluator-granularity-failure | (overall) | Whether your evaluator misses errors a finer-grained check would catch |
| intrinsic-contradiction | Input length | Whether contradictions rise with source length |
| entity-confusion | Content or document type | Whether entity-confusion is worse on certain document types |
Evaluator and annotation guidance
Three evaluator configurations cover the six families. entity-confusion uses programmatic matching; four families use rubric-human; the evaluator-granularity-failure family uses an LLM judge (because the point is to test whether the judge catches what the human catches).
Reference-answer cases (exact-or-programmatic)
Used for entity-confusion, where the reference captures the correct attribution.
Validity evidence needed:
- The reference answer captures the correct attribution or value, not a paraphrase that would force partial matching.
- A sample of the passes and failures is read by a person to confirm the comparison is judging what you think it is.
Metrics: pass rate with a confidence interval, per-slice pass rate.
Tools: pass rate CI calculator, slice and class balance analyzer.
Rubric cases (rubric-human)
Used for intrinsic-contradiction, extrinsic-addition, salient-omission, and granularity-mismatch.
Validity evidence needed:
- Two raters on a sample, with their agreement reported and corrected for chance.
- The rubric operates at the sentence or claim level, not the document level, for intrinsic and extrinsic error detection.
- Disagreements resolved by rewriting the rubric, not by averaging.
The sentence-level requirement is load-bearing. A document-level rubric (“Is this summary faithful? Yes/No”) misses the errors this pack is designed to catch. The rubric must ask about each claim in the summary against the passage it should trace to.
Metrics: inter-rater agreement, pass rate, per-claim support.
Tools: inter-rater reliability calculator, multi-rater agreement calculator.
LLM-judge cases (rubric, llm-judge)
Used for the evaluator-granularity-failure family, where the point is to test the judge itself.
Validity evidence needed:
- Agreement with human labels on your own items, measured and reported with an interval.
- The judge operates at the sentence or claim level. A document-level judge adds noise that a claim-level judge avoids.
- The judge is calibrated against known contradictions, known additions, and known omissions before it grades unknown cases.
Metrics: agreement with human labels, calibration, per-claim support.
Tools: judge validation report builder, inter-rater reliability calculator.
Starter rows
These rows are starter examples showing what a case in this family looks like. They are not validated gold-standard answers. Before any of them grades a real output, someone who knows your system needs to confirm or rewrite the expected behavior in each one.
Each row’s input names the source document in square brackets rather than carrying it. The bracket is a specification of the document shape the case needs; substitute a document of your own that matches it before running the row.
TABLEShow full table (10 rows)Showing full table (10 rows)
| ID | Input (abbreviated) | Case family | Oracle type | Evaluator | Reversibility | Blast radius |
|---|---|---|---|---|---|---|
| sum-001 | Summarize a quarterly earnings report in one paragraph (revenue up, net income down, new product launch in last paragraph) | salient-omission / intrinsic-contradiction | rubric | rubric-human | R1 | B1 |
| sum-002 | Summarize a clinical trial abstract (primary endpoint met, p=0.03, two serious adverse events) | salient-omission | rubric | rubric-human | R1 | B2 |
| sum-003 | Detailed summary of a services contract (liability cap in section 12, auto-renewal in section 15) | salient-omission | rubric | rubric-human | R1 | B1 |
| sum-004 | Compare two product specifications and summarize the differences | entity-confusion | rubric | rubric-human | R0 | B0 |
| sum-005 | Summarize a meeting transcript in bullet points (three action items assigned to named people) | entity-confusion / salient-omission | rubric | rubric-human | R0 | B1 |
| sum-006 | Summarize a merger news article in two sentences (acquirer X, target Y, price $4.2B) | entity-confusion | reference-answer | exact-or-programmatic | R1 | B1 |
| sum-007 | Summarize a research paper focusing on methodology (RCT, n=200, three arms) | extrinsic-addition / granularity-mismatch | rubric | rubric-human | R0 | B0 |
| sum-008 | Summarize three incident reports into one executive brief | entity-confusion | rubric | rubric-human | R0 | B1 |
| sum-009 | Summarize a safety data sheet (flash point, toxicity category, conditional first-aid instruction) | intrinsic-contradiction / salient-omission | rubric | rubric-human | R2 | B2 |
| sum-010 | Summarize section 8 only from a 12-section policy document | extrinsic-addition / granularity-mismatch | rubric | rubric-human | R0 | B0 |
Failure-to-case mapping
| Case family | Concern presets caught |
|---|---|
| intrinsic-contradiction | unsupported-claim |
| extrinsic-addition | unsupported-claim, omission |
| salient-omission | omission |
| granularity-mismatch | specification-violation, omission |
| entity-confusion | wrong-entity, unsupported-claim |
| evaluator-granularity-failure | verification-failure, unsupported-claim |
Use the risk-to-test mapper to trace from a named concern to the case families that test for it.
Downloads
Starter files for this pack, generated from the same source the examples above are read from: the CSV and the JSONL. Both carry the case profile in metadata.case. Every row is starting material, not a validated answer. This family’s files are published as summarization.csv and summarization.jsonl. To edit them in the browser and add your own, use the blueprint builder.
Which instruments apply
- Pass rate CI calculator for a confidence interval on the pass rate per family.
- Slice and class balance analyzer for checking whether the set covers each case-by-slice cell.
- Inter-rater reliability calculator for the agreement check on rubric-graded families.
- Multi-rater agreement calculator for agreement across more than two raters.
- Judge validation report builder for calibrating and validating the LLM judge against human labels.
- Eval coverage matrix for the grid of which cells are populated and which are empty.
- Golden set blueprint builder for editing and extending the starter rows.
- Risk-to-test mapper for mapping from a named concern to the case families that test for it.
Companion pack guides
- Structured extraction tests a related fidelity surface: whether the system transfers content from a source into a structured output without adding or dropping fields.
- Customer support covers multi-turn conversation, where the failure families include complaint recognition and escalation handoff alongside policy compliance.