LatentEval

For builders

Golden set blueprint: RAG and knowledge assistants

Seven case families for a RAG assistant, five slice families, oracle types and evaluator guidance per family, downloadable starter rows, and a pharma overlay.

For builders

In brief

3 POINTS
  • One hallucination rate rolls seven different failure points into a single number you cannot act on.
  • You score retrieval and generation on separate denominators, because perfect retrieval is compatible with an answer that never used the retrieved passage.
  • Stale material is its own case class: the document was correct when it was indexed and is no longer correct now.

Why this family’s correctness is different

A single hallucination rate collapses seven distinct failure points into one number nobody can act on: the answer was not in the corpus, the passage was not retrieved, the passage was retrieved and ignored, the source was stale, the synthesis was incomplete, the specificity was wrong, or the format was wrong. Score retrieval and generation separately with an interval on each. A perfect retrieval score is compatible with an answer that ignored the passage, so a generation check exists whether or not retrieval was measured.

This pack builds on research published on this site. A taxonomy of RAG pipeline failure modes identifies the seven failure points the case families below are drawn from. A practical guide to evaluating RAG pipelines establishes the retrieval-versus-generation scoring split. And the silent failure problem in AI evaluations documents how a single accuracy number hides the failure shape.

Case families

Each case family tests one failure point. The first three sit on the retrieval side of the pipeline, the next four on the generation side. A set that covers only one side reports nothing about the other.

These counts seed a starter set; none of them supports a rate claim on its own. The coverage matrix floors set the claimable and readable thresholds.

A set built from this pack carries several oracle types across its families; run the evaluator navigator once per family and write acceptance criteria per family.

TABLEShow full table (7 rows)Showing full table (7 rows)
Case familyWhat it testsExpected behaviorUnacceptable behaviorOracle typeEvaluatorReviewer expertiseSeverity treatmentStarter rows to writeWhat this count licenses
answer-not-in-corpusWhether the system says it does not know rather than constructing a plausible answer from nothingA statement that the available material does not cover this, with no fabricated answerGenerates a fluent answer from parametric memory; returns a topically adjacent passage that does not answer the questionrubricrubric-humanSomeone who knows the corpus well enough to confirm the answer genuinely cannot be foundSeverity follows the action the reader would take on the fabricated answer15 - 30At 30, worst-case 95% interval: +/-16.8 points
retrievable-not-rankedWhether the relevant passage reaches the generation context at allThe passage is ranked within the top-k and included in the generation contextThe passage is in the index and does not appear in the top-k; a near-miss passage ranks above itreference-answerexact-or-programmatic(none)Retrieval failures are counted against their own denominator, separate from generation20 - 40At 40, worst-case 95% interval: +/-14.8 points
retrieved-not-usedWhether the generator uses the retrieved passage when it is presentThe answer draws on the retrieved passage and cites or reflects its contentAnswers from parametric memory when the passage was in context; cites the passage and says something unsupportedrubricrubric-humanSomeone who can read the passage and the answer side by sideSeverity follows the difference between what the passage says and what the answer says15 - 25At 25, worst-case 95% interval: +/-18.2 points
stale-sourceWhether the system answers from a superseded version of a document when a current one existsThe answer uses the current version, or says the answer may have changed and names the version it usedAnswers from the superseded version without noting it; ranks the outdated version above the current onereference-answerexact-or-programmatic(none)A stale answer about a policy someone acts on inherits the blast radius of that action10 - 20At 20, worst-case 95% interval: +/-20.1 points
multi-passage-synthesisWhether the system can synthesize information from two or more retrieved passagesAn answer that draws on all the relevant passages and does not contradict any of themAnswers from the first passage only; invents a reconciliation between passages that contradict each otherrubricrubric-humanSomeone who can read both passages and judge fidelityAn incomplete synthesis is an omission; an overstated reconciliation is an unsupported claim10 - 20At 20, worst-case 95% interval: +/-20.1 points
wrong-specificityWhether the answer gives the specific value the question asked for rather than a general statementThe specific value the question asked for, drawn from the passageA general statement when a specific value was asked for; a specific value when an overview was asked forreference-answerexact-or-programmatic(none)Severity follows what the reader would do with the wrong level of detail10 - 20At 20, worst-case 95% interval: +/-20.1 points
format-mismatchWhether the output matches the requested format when a format was specifiedThe content is correct and the format matches what was asked forCorrect content in prose when a table was requested; a table with correct values and missing column headersrubricrubric-human(none)Usually low severity unless the downstream consumer is a parser that will reject the format8 - 15At 15, worst-case 95% interval: +/-22.6 points

The seven families split into a retrieval stage (answer-not-in-corpus, retrievable-not-ranked, stale-source) and a generation stage (retrieved-not-used, multi-passage-synthesis, wrong-specificity, format-mismatch). A hallucination rate that rolls all seven into one number cannot point you at the stage where the problem sits. Score retrieval and generation with their own denominators.

Slice families

Slices subdivide the case set by a dimension where failure rates are expected to differ. Each slice runs across every case family.

TABLEShow full table (5 rows)Showing full table (5 rows)
Slice familyWhy it mattersExample valuesSensitive
Content or document typeA policy PDF, an FAQ page, and a safety data sheet carry different structures and different failure shapespolicy document, FAQ, technical specification, safety data sheetNo
Input lengthA one-line question and a multi-paragraph request with embedded constraints fail in different waysunder 50 words, 50 to 200 words, over 200 wordsNo
Language and localeRetrieval quality often differs by language, and a locale-specific answer drawn from a global corpus is a different failure from a wrong answeren-US, de-DE, ja-JP, mixed-language queryNo
Time or recencyA question about the current policy and a question about last year’s policy require different retrieval, and a stale answer to the first is a failure the second does not havecurrent version only, historical version, time-comparativeNo
Upstream sourceA typed question and a voice transcript carry different noise levels, and retrieval degrades differently on eachtyped by a person, voice transcript, forwarded email, another agentNo

None of these slices is marked sensitive. Where a slice touches a population or a jurisdiction, the sensitive flag would change the storage and display rules for the data in that cell.

Coverage matrix skeleton

The coverage matrix is a grid of case families against slice families. Each cell below names what you would not know if the cell were empty. Use the eval coverage matrix tool to fill it against your own system.

TABLEShow full table (11 rows)Showing full table (11 rows)
Case familySlice familyWhat an empty cell hides
answer-not-in-corpusContent or document typeWhether the system fabricates answers differently by document type
answer-not-in-corpus(overall)Whether the system says it does not know or fabricates an answer
retrievable-not-rankedContent or document typeWhether retrieval quality varies by document type
retrievable-not-rankedLanguage and localeWhether the retriever works as well in other languages
retrieved-not-usedInput lengthWhether the generator ignores passages more often with longer queries
retrieved-not-used(overall)Whether the generator uses the passages it receives
stale-sourceTime or recencyWhether the system surfaces outdated material
multi-passage-synthesisContent or document typeWhether synthesis fails more often across document types
wrong-specificity(overall)Whether the system gives a general answer when a specific one was asked for
format-mismatch(overall)Whether the system produces the format the user asked for
retrievable-not-rankedUpstream sourceWhether retrieval degrades on voice transcripts or forwarded email

Evaluator and annotation guidance

Each oracle type admits a bounded set of evaluators. The evaluator needs validity evidence before its numbers mean anything.

Reference-answer cases (exact-or-programmatic)

Used for retrievable-not-ranked, stale-source, and wrong-specificity families, where a ground-truth value exists.

Validity evidence needed:

  • The normalization is written down before anything is graded: case folding, whitespace, number formats, and unit equivalences.
  • A sample of the passes and the failures is read by a person once, to confirm the comparison is judging what you think it is.

Metrics: pass rate with a confidence interval, per-slice pass rate, retrieval rank.

Tools: pass rate CI calculator, RAG retrieval metrics calculator, slice and class balance analyzer.

Rubric cases (rubric-human)

Used for answer-not-in-corpus, retrieved-not-used, multi-passage-synthesis, and format-mismatch, where the answer needs a human judgment against a written rubric.

Validity evidence needed:

  • Two raters on a sample, with their agreement reported and corrected for chance.
  • A written rubric that a second rater can apply without asking the first one what it meant.
  • Disagreements resolved by rewriting the rubric, not by averaging the two verdicts.

Metrics: inter-rater agreement, pass rate, per-claim support.

Tools: inter-rater reliability calculator, multi-rater agreement calculator.

Expected-behavior-only cases (human-expert)

Used for overlay families where correct behavior has not been turned into a checkable oracle.

Validity evidence needed:

  • The expert who wrote the expected behavior is named and their domain qualification is stated.
  • The row is not gradeable until the expert has validated it.

Metrics: constraint violation count.

Tools: pass rate CI calculator.

Starter rows

These rows are starter examples showing what a case in this family looks like. They are not validated gold-standard answers. Before any of them grades a real output, someone who knows your system needs to confirm or rewrite the expected behavior in each one.

TABLEShow full table (10 rows)Showing full table (10 rows)
IDInput (abbreviated)Case familyOracle typeEvaluatorReversibilityBlast radius
rag-001Maximum parcel weight for express shipping to rural addresses?wrong-specificityreference-answerexact-or-programmaticR0B0
rag-002Company policy on using generative AI for client deliverables?answer-not-in-corpusrubricrubric-humanR0B1
rag-003Key changes between the 2025 and 2026 employee handbook editionsmulti-passage-synthesisrubricrubric-humanR0B1
rag-004Return window before the April 2026 policy update?stale-sourcereference-answerexact-or-programmaticR0B1
rag-005Steps to request a building access card, in orderformat-mismatchrubricrubric-humanR0B0
rag-006Oslo office: local requirements for employee data retentionretrieved-not-usedrubricrubric-humanR1B1
rag-007Vacation days for part-time employees after five years?wrong-specificityreference-answerexact-or-programmaticR0B0
rag-008Is the SDS flash point for product X-140 still 65C?stale-sourcerubricrubric-humanR1B2
rag-009Turn 1: APAC expense policy. Turn 2: What is the meal per diem?retrieved-not-usedrubricrubric-humanR0B0
rag-010Incident response procedure as a table with four columnsformat-mismatchrubricrubric-humanR0B0

Failure-to-case mapping

Each case family catches one or more of the named concerns in the risk-to-test mapper. The mapping below shows which concerns each family is designed to surface.

Case familyConcern presets caught
answer-not-in-corpusunsupported-claim, over-refusal
retrievable-not-rankedstale-retrieval, omission
retrieved-not-usedunsupported-claim, omission
stale-sourcestale-retrieval, unsupported-claim
multi-passage-synthesisomission, unsupported-claim
wrong-specificityomission, unsupported-claim
format-mismatchformat-contract-violation

Downloads

Starter files for this pack, generated from the same source the examples above are read from: the CSV and the JSONL. Both carry the case profile in metadata.case. Every row is starting material, not a validated answer. This family’s files are published as knowledge-rag.csv and knowledge-rag.jsonl. To edit them in the browser and add your own, use the blueprint builder.

Pharma medical-information overlay

This section lists case classes that practitioners in this area commonly test for. It is not legal, regulatory, medical or compliance advice and does not establish what is required of you. Passing these cases does not make a system compliant with anything. Take them to your own regulatory, legal or compliance function and let that function decide what correct behavior is before you grade against it.

Every row in this section carries an expected behavior that a domain expert must validate before it grades anything. In the downloadable files these rows are tagged provenance=overlay-starter-needs-expert-validation, so you can filter or remove them in one pass.

This overlay ships because it changes four things the base pack does not cover:

  • Case taxonomy. Solicited-versus-unsolicited request discrimination, on-label versus off-label routing, and citation fidelity to a medical source are case classes the base RAG pack has no reason to carry.
  • Evaluator. Every overlay family is expected-behavior-only, graded by a human expert, because the correct behavior has not been turned into a checkable oracle.
  • Expertise. The reviewer is in the medical-information, medical-affairs, or pharmacovigilance function, not a general knowledge owner.
  • Severity treatment. A wrong answer about a medical product carries B3 blast radius by default because the reader may act on it outside the organization.

Overlay case families

TABLEShow full table (3 rows)Showing full table (3 rows)
Case familyWhat it testsEvaluatorReviewer expertiseStarter rows to writeWhat this count licenses
solicited-vs-unsolicitedWhether the system correctly discriminates between a solicited medical-information request and an unsolicited onehuman-expertMedical-information or medical-affairs function8 - 15At 15, worst-case 95% interval: +/-22.6 points
off-label-routingWhether the system routes an off-label question to the appropriate channel rather than answering it directlyhuman-expertMedical affairs or regulatory affairs6 - 12At 12, worst-case 95% interval: +/-24.6 points
citation-fidelityWhether a cited source supports the claim made about it, with no omitted safety findings and no overstated efficacyhuman-expertMedical affairs or pharmacovigilance6 - 12At 12, worst-case 95% interval: +/-24.6 points

Overlay starter rows

TABLEShow full table (4 rows)Showing full table (4 rows)
IDInput (abbreviated)Overlay familyReversibilityBlast radius
rag-pharma-001Physician asks about efficacy of drug A for a condition not on the approved labeloff-label-routingR2B3
rag-pharma-002Nurse practitioner asks whether drug B can be used for weight management (approved for diabetes)solicited-vs-unsolicitedR2B3
rag-pharma-003Key safety findings from phase III trial of drug C, citing the published sourcecitation-fidelityR2B3
rag-pharma-004Patient calls about side effects after starting drug D, describing symptoms consistent with a known adverse eventadverse-event-detection (citation-fidelity family)R2B3

Every overlay row requires domain-expert validation before it grades anything. The pass logic for each row is written by the domain expert who validates the case. Until then, the row is not gradeable, and running an automated grader over it produces a number about a rule nobody wrote.

Which instruments apply

The following live tools support the measurements this pack describes. Each links to a tool that accepts input from a golden set built on these case families.

Companion pack guides

  • Summarization tests a closely related failure surface: whether compressed output preserves what matters, with separate families for contradictions, unsupported additions, and omissions.
  • Structured extraction covers field-level accuracy, where the shared concern is whether the system faithfully transfers content from a source document into a structured output.