LatentEval

For builders

Golden set blueprint: summarization

Six case families for summarization: source contradictions, unsupported additions, and salient omissions counted separately, with evaluator granularity tested as a variable.

For builders

In brief

3 POINTS
  • A contradiction and an unsupported addition are different failures, and the distinction holds only once you name a source for each claim.
  • Evaluator granularity is itself a variable in the result: a document-level check misses errors that a sentence-level check catches.
  • You score omissions against what the summary was asked to preserve, not against the full length of the source.

Why this family’s correctness is different

A faithfulness number that does not say whether the error was a contradiction, an unsupported addition, or a dropped finding is a number that cannot point you at the fix. Intrinsic errors contradict the source; extrinsic errors add something the source never said; omissions drop something the source emphasized. The distinction only means anything once a source is named for each claim. And the evaluator granularity matters: a document-level consistency check misses errors a sentence-level check catches, so the granularity of the evaluator is itself a variable in the result.

This pack builds on research published on this site. The silent failure problem in AI evaluations documents how a single faithfulness number hides the failure shape. How LLM-as-a-judge evaluations work and where they break covers the granularity problem: a judge applied at the wrong level of analysis adds noise that looks like signal. A practical framework for AI agent evaluation establishes the general principle that each failure type needs its own denominator.

Case families

Six families. The first three (intrinsic-contradiction, extrinsic-addition, salient-omission) separate what a single faithfulness number collapses. The fourth (granularity-mismatch) catches the case where the summary does not match the requested detail level. The fifth (entity-confusion) catches the case where the summary attributes a fact to the wrong entity. The sixth (evaluator-granularity-failure) is a meta-case: it tests whether the chosen evaluator operates at a fine enough granularity to catch the errors the other families describe.

These counts seed a starter set; none of them supports a rate claim on its own. The coverage matrix floors set the claimable and readable thresholds.

A set built from this pack carries several oracle types across its families; run the evaluator navigator once per family and write acceptance criteria per family.

TABLEShow full table (6 rows)Showing full table (6 rows)
Case familyWhat it testsExpected behaviorUnacceptable behaviorOracle typeEvaluatorReviewer expertiseSeverity treatmentStarter rows to writeWhat this count licenses
intrinsic-contradictionWhether the summary states something the source explicitly says differentlyEvery claim is consistent with the corresponding passage in the sourceStates a number, date, or name that differs; reverses a causal direction the source statesrubricrubric-humanSomeone who can spot a factual conflict between the summary and the sourceSeverity follows the claim. Count contradictions separately from extrinsic errors15 - 30At 30, worst-case 95% interval: +/-16.8 points
extrinsic-additionWhether the summary introduces a claim the source is silent on, true or notEvery claim traces to a passage in the source or is explicitly marked as externalStates a fact the source does not contain without flagging it; infers a conclusion and presents it as the source’s findingrubricrubric-humanSomeone who can confirm whether a claim is supported, not just plausibleAn extrinsic-addition that happens to be true in the world is still an error against the source12 - 25At 25, worst-case 95% interval: +/-18.2 points
salient-omissionWhether the summary preserves the points the source emphasizesEvery conclusion, finding, warning, or recommendation the source calls out appearsDrops a safety warning; covers positive findings and omits the limitationsrubricrubric-humanSomeone who can decide which points are salient, not just presentCount omissions against a stated preservation target, not an abstract completeness goal12 - 25At 25, worst-case 95% interval: +/-18.2 points
granularity-mismatchWhether the summary matches the level of detail the request asks forDetail level matches: a one-paragraph summary does not itemize every findingA detailed summary when one paragraph was asked for; three sentences when comprehensive was asked forrubricrubric-humanSomeone who can judge whether the detail level serves the stated purposeUsually low severity on its own. Rises when the wrong granularity causes an omission8 - 15At 15, worst-case 95% interval: +/-22.6 points
entity-confusionWhether the summary attributes a fact to the right entity when the source discusses severalEvery attribution matches the entity the source assigns it toAttributes company A’s figure to company B; merges two people’s statements into onereference-answerexact-or-programmatic(none)A misattributed figure carries the severity of the decision it feeds8 - 15At 15, worst-case 95% interval: +/-22.6 points
evaluator-granularity-failureWhether the chosen evaluator operates at a granularity fine enough to detect the error classThe evaluation catches the seeded error at the sentence or claim levelA document-level check passes a summary that contradicts one sentence; NLI at full-document level misses sentence-level inconsistencyrubricllm-judgeSomeone who can compare sentence-level and document-level verdictsA passed check that should have failed undermines the numbers for every other case6 - 12At 12, worst-case 95% interval: +/-24.6 points

The evaluator-granularity-failure family is a meta-case. It does not test the summarizer; it tests the evaluator. A document-level faithfulness check that passes a summary containing a sentence-level contradiction has not measured faithfulness at the resolution that matters. Including these cases in the set means the set can detect when its own evaluation is too coarse.

Slice families

TABLEShow full table (5 rows)Showing full table (5 rows)
Slice familyWhy it mattersExample valuesSensitive
Content or document typeA clinical abstract, a legal contract, and a meeting transcript carry different structures and error shapesclinical abstract, legal contract, earnings report, meeting transcript, policy documentNo
Input lengthOmissions rise with source length. A set built on short documents says nothing about a system that summarizes long onesunder 1,000 words, 1,000 to 5,000 words, over 5,000 wordsNo
Intent or task typeAn executive summary, a comparative summary, and a safety-critical summary have different preservation targets and failure costsexecutive summary, detailed summary, comparative, section-scoped, action-item extractionNo
Language and localeSummarization quality often differs by language, and a cross-language summary adds a translation failure surfaceen-US, de-DE, ja-JP, cross-languageNo
Upstream sourceAn OCR-scanned PDF and a clean digital document carry different noise levels, and extraction errors in the source propagate into the summaryclean digital document, OCR scan, speech-to-text transcript, web scrapeNo

The input-length slice deserves special attention. Omissions are the error class most sensitive to source length: a system that omits nothing from a 500-word abstract may omit material findings from a 5,000-word report. A set built entirely on short sources has a structural blind spot.

Coverage matrix skeleton

Use the eval coverage matrix tool to fill the grid against your own system.

TABLEShow full table (11 rows)Showing full table (11 rows)
Case familySlice familyWhat an empty cell hides
intrinsic-contradictionContent or document typeWhether contradiction rates differ by document type
intrinsic-contradiction(overall)Whether the summarizer contradicts the source at all
extrinsic-additionContent or document typeWhether the summarizer fabricates more on one document type than another
extrinsic-addition(overall)Whether the summarizer adds unsupported claims
salient-omissionInput lengthWhether omissions rise with source length
salient-omission(overall)Whether the summarizer drops important points
granularity-mismatchIntent or task typeWhether the system matches the requested detail level
entity-confusion(overall)Whether the summarizer swaps entities when the source discusses several
evaluator-granularity-failure(overall)Whether your evaluator misses errors a finer-grained check would catch
intrinsic-contradictionInput lengthWhether contradictions rise with source length
entity-confusionContent or document typeWhether entity-confusion is worse on certain document types

Evaluator and annotation guidance

Three evaluator configurations cover the six families. entity-confusion uses programmatic matching; four families use rubric-human; the evaluator-granularity-failure family uses an LLM judge (because the point is to test whether the judge catches what the human catches).

Reference-answer cases (exact-or-programmatic)

Used for entity-confusion, where the reference captures the correct attribution.

Validity evidence needed:

  • The reference answer captures the correct attribution or value, not a paraphrase that would force partial matching.
  • A sample of the passes and failures is read by a person to confirm the comparison is judging what you think it is.

Metrics: pass rate with a confidence interval, per-slice pass rate.

Tools: pass rate CI calculator, slice and class balance analyzer.

Rubric cases (rubric-human)

Used for intrinsic-contradiction, extrinsic-addition, salient-omission, and granularity-mismatch.

Validity evidence needed:

  • Two raters on a sample, with their agreement reported and corrected for chance.
  • The rubric operates at the sentence or claim level, not the document level, for intrinsic and extrinsic error detection.
  • Disagreements resolved by rewriting the rubric, not by averaging.

The sentence-level requirement is load-bearing. A document-level rubric (“Is this summary faithful? Yes/No”) misses the errors this pack is designed to catch. The rubric must ask about each claim in the summary against the passage it should trace to.

Metrics: inter-rater agreement, pass rate, per-claim support.

Tools: inter-rater reliability calculator, multi-rater agreement calculator.

LLM-judge cases (rubric, llm-judge)

Used for the evaluator-granularity-failure family, where the point is to test the judge itself.

Validity evidence needed:

  • Agreement with human labels on your own items, measured and reported with an interval.
  • The judge operates at the sentence or claim level. A document-level judge adds noise that a claim-level judge avoids.
  • The judge is calibrated against known contradictions, known additions, and known omissions before it grades unknown cases.

Metrics: agreement with human labels, calibration, per-claim support.

Tools: judge validation report builder, inter-rater reliability calculator.

Starter rows

These rows are starter examples showing what a case in this family looks like. They are not validated gold-standard answers. Before any of them grades a real output, someone who knows your system needs to confirm or rewrite the expected behavior in each one.

Each row’s input names the source document in square brackets rather than carrying it. The bracket is a specification of the document shape the case needs; substitute a document of your own that matches it before running the row.

TABLEShow full table (10 rows)Showing full table (10 rows)
IDInput (abbreviated)Case familyOracle typeEvaluatorReversibilityBlast radius
sum-001Summarize a quarterly earnings report in one paragraph (revenue up, net income down, new product launch in last paragraph)salient-omission / intrinsic-contradictionrubricrubric-humanR1B1
sum-002Summarize a clinical trial abstract (primary endpoint met, p=0.03, two serious adverse events)salient-omissionrubricrubric-humanR1B2
sum-003Detailed summary of a services contract (liability cap in section 12, auto-renewal in section 15)salient-omissionrubricrubric-humanR1B1
sum-004Compare two product specifications and summarize the differencesentity-confusionrubricrubric-humanR0B0
sum-005Summarize a meeting transcript in bullet points (three action items assigned to named people)entity-confusion / salient-omissionrubricrubric-humanR0B1
sum-006Summarize a merger news article in two sentences (acquirer X, target Y, price $4.2B)entity-confusionreference-answerexact-or-programmaticR1B1
sum-007Summarize a research paper focusing on methodology (RCT, n=200, three arms)extrinsic-addition / granularity-mismatchrubricrubric-humanR0B0
sum-008Summarize three incident reports into one executive briefentity-confusionrubricrubric-humanR0B1
sum-009Summarize a safety data sheet (flash point, toxicity category, conditional first-aid instruction)intrinsic-contradiction / salient-omissionrubricrubric-humanR2B2
sum-010Summarize section 8 only from a 12-section policy documentextrinsic-addition / granularity-mismatchrubricrubric-humanR0B0

Failure-to-case mapping

Case familyConcern presets caught
intrinsic-contradictionunsupported-claim
extrinsic-additionunsupported-claim, omission
salient-omissionomission
granularity-mismatchspecification-violation, omission
entity-confusionwrong-entity, unsupported-claim
evaluator-granularity-failureverification-failure, unsupported-claim

Use the risk-to-test mapper to trace from a named concern to the case families that test for it.

Downloads

Starter files for this pack, generated from the same source the examples above are read from: the CSV and the JSONL. Both carry the case profile in metadata.case. Every row is starting material, not a validated answer. This family’s files are published as summarization.csv and summarization.jsonl. To edit them in the browser and add your own, use the blueprint builder.

Which instruments apply

Companion pack guides

  • Structured extraction tests a related fidelity surface: whether the system transfers content from a source into a structured output without adding or dropping fields.
  • Customer support covers multi-turn conversation, where the failure families include complaint recognition and escalation handoff alongside policy compliance.