For builders
Golden set blueprint: RAG and knowledge assistants
Seven case families for a RAG assistant, five slice families, oracle types and evaluator guidance per family, downloadable starter rows, and a pharma overlay.
In brief
3 POINTS- One hallucination rate rolls seven different failure points into a single number you cannot act on.
- You score retrieval and generation on separate denominators, because perfect retrieval is compatible with an answer that never used the retrieved passage.
- Stale material is its own case class: the document was correct when it was indexed and is no longer correct now.
On this page (11)
Why this family’s correctness is different
A single hallucination rate collapses seven distinct failure points into one number nobody can act on: the answer was not in the corpus, the passage was not retrieved, the passage was retrieved and ignored, the source was stale, the synthesis was incomplete, the specificity was wrong, or the format was wrong. Score retrieval and generation separately with an interval on each. A perfect retrieval score is compatible with an answer that ignored the passage, so a generation check exists whether or not retrieval was measured.
This pack builds on research published on this site. A taxonomy of RAG pipeline failure modes identifies the seven failure points the case families below are drawn from. A practical guide to evaluating RAG pipelines establishes the retrieval-versus-generation scoring split. And the silent failure problem in AI evaluations documents how a single accuracy number hides the failure shape.
Case families
Each case family tests one failure point. The first three sit on the retrieval side of the pipeline, the next four on the generation side. A set that covers only one side reports nothing about the other.
These counts seed a starter set; none of them supports a rate claim on its own. The coverage matrix floors set the claimable and readable thresholds.
A set built from this pack carries several oracle types across its families; run the evaluator navigator once per family and write acceptance criteria per family.
TABLEShow full table (7 rows)Showing full table (7 rows)
| Case family | What it tests | Expected behavior | Unacceptable behavior | Oracle type | Evaluator | Reviewer expertise | Severity treatment | Starter rows to write | What this count licenses |
|---|---|---|---|---|---|---|---|---|---|
| answer-not-in-corpus | Whether the system says it does not know rather than constructing a plausible answer from nothing | A statement that the available material does not cover this, with no fabricated answer | Generates a fluent answer from parametric memory; returns a topically adjacent passage that does not answer the question | rubric | rubric-human | Someone who knows the corpus well enough to confirm the answer genuinely cannot be found | Severity follows the action the reader would take on the fabricated answer | 15 - 30 | At 30, worst-case 95% interval: +/-16.8 points |
| retrievable-not-ranked | Whether the relevant passage reaches the generation context at all | The passage is ranked within the top-k and included in the generation context | The passage is in the index and does not appear in the top-k; a near-miss passage ranks above it | reference-answer | exact-or-programmatic | (none) | Retrieval failures are counted against their own denominator, separate from generation | 20 - 40 | At 40, worst-case 95% interval: +/-14.8 points |
| retrieved-not-used | Whether the generator uses the retrieved passage when it is present | The answer draws on the retrieved passage and cites or reflects its content | Answers from parametric memory when the passage was in context; cites the passage and says something unsupported | rubric | rubric-human | Someone who can read the passage and the answer side by side | Severity follows the difference between what the passage says and what the answer says | 15 - 25 | At 25, worst-case 95% interval: +/-18.2 points |
| stale-source | Whether the system answers from a superseded version of a document when a current one exists | The answer uses the current version, or says the answer may have changed and names the version it used | Answers from the superseded version without noting it; ranks the outdated version above the current one | reference-answer | exact-or-programmatic | (none) | A stale answer about a policy someone acts on inherits the blast radius of that action | 10 - 20 | At 20, worst-case 95% interval: +/-20.1 points |
| multi-passage-synthesis | Whether the system can synthesize information from two or more retrieved passages | An answer that draws on all the relevant passages and does not contradict any of them | Answers from the first passage only; invents a reconciliation between passages that contradict each other | rubric | rubric-human | Someone who can read both passages and judge fidelity | An incomplete synthesis is an omission; an overstated reconciliation is an unsupported claim | 10 - 20 | At 20, worst-case 95% interval: +/-20.1 points |
| wrong-specificity | Whether the answer gives the specific value the question asked for rather than a general statement | The specific value the question asked for, drawn from the passage | A general statement when a specific value was asked for; a specific value when an overview was asked for | reference-answer | exact-or-programmatic | (none) | Severity follows what the reader would do with the wrong level of detail | 10 - 20 | At 20, worst-case 95% interval: +/-20.1 points |
| format-mismatch | Whether the output matches the requested format when a format was specified | The content is correct and the format matches what was asked for | Correct content in prose when a table was requested; a table with correct values and missing column headers | rubric | rubric-human | (none) | Usually low severity unless the downstream consumer is a parser that will reject the format | 8 - 15 | At 15, worst-case 95% interval: +/-22.6 points |
The seven families split into a retrieval stage (answer-not-in-corpus, retrievable-not-ranked, stale-source) and a generation stage (retrieved-not-used, multi-passage-synthesis, wrong-specificity, format-mismatch). A hallucination rate that rolls all seven into one number cannot point you at the stage where the problem sits. Score retrieval and generation with their own denominators.
Slice families
Slices subdivide the case set by a dimension where failure rates are expected to differ. Each slice runs across every case family.
TABLEShow full table (5 rows)Showing full table (5 rows)
| Slice family | Why it matters | Example values | Sensitive |
|---|---|---|---|
| Content or document type | A policy PDF, an FAQ page, and a safety data sheet carry different structures and different failure shapes | policy document, FAQ, technical specification, safety data sheet | No |
| Input length | A one-line question and a multi-paragraph request with embedded constraints fail in different ways | under 50 words, 50 to 200 words, over 200 words | No |
| Language and locale | Retrieval quality often differs by language, and a locale-specific answer drawn from a global corpus is a different failure from a wrong answer | en-US, de-DE, ja-JP, mixed-language query | No |
| Time or recency | A question about the current policy and a question about last year’s policy require different retrieval, and a stale answer to the first is a failure the second does not have | current version only, historical version, time-comparative | No |
| Upstream source | A typed question and a voice transcript carry different noise levels, and retrieval degrades differently on each | typed by a person, voice transcript, forwarded email, another agent | No |
None of these slices is marked sensitive. Where a slice touches a population or a jurisdiction, the sensitive flag would change the storage and display rules for the data in that cell.
Coverage matrix skeleton
The coverage matrix is a grid of case families against slice families. Each cell below names what you would not know if the cell were empty. Use the eval coverage matrix tool to fill it against your own system.
TABLEShow full table (11 rows)Showing full table (11 rows)
| Case family | Slice family | What an empty cell hides |
|---|---|---|
| answer-not-in-corpus | Content or document type | Whether the system fabricates answers differently by document type |
| answer-not-in-corpus | (overall) | Whether the system says it does not know or fabricates an answer |
| retrievable-not-ranked | Content or document type | Whether retrieval quality varies by document type |
| retrievable-not-ranked | Language and locale | Whether the retriever works as well in other languages |
| retrieved-not-used | Input length | Whether the generator ignores passages more often with longer queries |
| retrieved-not-used | (overall) | Whether the generator uses the passages it receives |
| stale-source | Time or recency | Whether the system surfaces outdated material |
| multi-passage-synthesis | Content or document type | Whether synthesis fails more often across document types |
| wrong-specificity | (overall) | Whether the system gives a general answer when a specific one was asked for |
| format-mismatch | (overall) | Whether the system produces the format the user asked for |
| retrievable-not-ranked | Upstream source | Whether retrieval degrades on voice transcripts or forwarded email |
Evaluator and annotation guidance
Each oracle type admits a bounded set of evaluators. The evaluator needs validity evidence before its numbers mean anything.
Reference-answer cases (exact-or-programmatic)
Used for retrievable-not-ranked, stale-source, and wrong-specificity families, where a ground-truth value exists.
Validity evidence needed:
- The normalization is written down before anything is graded: case folding, whitespace, number formats, and unit equivalences.
- A sample of the passes and the failures is read by a person once, to confirm the comparison is judging what you think it is.
Metrics: pass rate with a confidence interval, per-slice pass rate, retrieval rank.
Tools: pass rate CI calculator, RAG retrieval metrics calculator, slice and class balance analyzer.
Rubric cases (rubric-human)
Used for answer-not-in-corpus, retrieved-not-used, multi-passage-synthesis, and format-mismatch, where the answer needs a human judgment against a written rubric.
Validity evidence needed:
- Two raters on a sample, with their agreement reported and corrected for chance.
- A written rubric that a second rater can apply without asking the first one what it meant.
- Disagreements resolved by rewriting the rubric, not by averaging the two verdicts.
Metrics: inter-rater agreement, pass rate, per-claim support.
Tools: inter-rater reliability calculator, multi-rater agreement calculator.
Expected-behavior-only cases (human-expert)
Used for overlay families where correct behavior has not been turned into a checkable oracle.
Validity evidence needed:
- The expert who wrote the expected behavior is named and their domain qualification is stated.
- The row is not gradeable until the expert has validated it.
Metrics: constraint violation count.
Tools: pass rate CI calculator.
Starter rows
These rows are starter examples showing what a case in this family looks like. They are not validated gold-standard answers. Before any of them grades a real output, someone who knows your system needs to confirm or rewrite the expected behavior in each one.
TABLEShow full table (10 rows)Showing full table (10 rows)
| ID | Input (abbreviated) | Case family | Oracle type | Evaluator | Reversibility | Blast radius |
|---|---|---|---|---|---|---|
| rag-001 | Maximum parcel weight for express shipping to rural addresses? | wrong-specificity | reference-answer | exact-or-programmatic | R0 | B0 |
| rag-002 | Company policy on using generative AI for client deliverables? | answer-not-in-corpus | rubric | rubric-human | R0 | B1 |
| rag-003 | Key changes between the 2025 and 2026 employee handbook editions | multi-passage-synthesis | rubric | rubric-human | R0 | B1 |
| rag-004 | Return window before the April 2026 policy update? | stale-source | reference-answer | exact-or-programmatic | R0 | B1 |
| rag-005 | Steps to request a building access card, in order | format-mismatch | rubric | rubric-human | R0 | B0 |
| rag-006 | Oslo office: local requirements for employee data retention | retrieved-not-used | rubric | rubric-human | R1 | B1 |
| rag-007 | Vacation days for part-time employees after five years? | wrong-specificity | reference-answer | exact-or-programmatic | R0 | B0 |
| rag-008 | Is the SDS flash point for product X-140 still 65C? | stale-source | rubric | rubric-human | R1 | B2 |
| rag-009 | Turn 1: APAC expense policy. Turn 2: What is the meal per diem? | retrieved-not-used | rubric | rubric-human | R0 | B0 |
| rag-010 | Incident response procedure as a table with four columns | format-mismatch | rubric | rubric-human | R0 | B0 |
Failure-to-case mapping
Each case family catches one or more of the named concerns in the risk-to-test mapper. The mapping below shows which concerns each family is designed to surface.
| Case family | Concern presets caught |
|---|---|
| answer-not-in-corpus | unsupported-claim, over-refusal |
| retrievable-not-ranked | stale-retrieval, omission |
| retrieved-not-used | unsupported-claim, omission |
| stale-source | stale-retrieval, unsupported-claim |
| multi-passage-synthesis | omission, unsupported-claim |
| wrong-specificity | omission, unsupported-claim |
| format-mismatch | format-contract-violation |
Downloads
Starter files for this pack, generated from the same source the examples above are read from: the CSV and the JSONL. Both carry the case profile in metadata.case. Every row is starting material, not a validated answer. This family’s files are published as knowledge-rag.csv and knowledge-rag.jsonl. To edit them in the browser and add your own, use the blueprint builder.
Pharma medical-information overlay
This section lists case classes that practitioners in this area commonly test for. It is not legal, regulatory, medical or compliance advice and does not establish what is required of you. Passing these cases does not make a system compliant with anything. Take them to your own regulatory, legal or compliance function and let that function decide what correct behavior is before you grade against it.
Every row in this section carries an expected behavior that a domain expert must validate before it grades anything. In the downloadable files these rows are tagged provenance=overlay-starter-needs-expert-validation, so you can filter or remove them in one pass.
This overlay ships because it changes four things the base pack does not cover:
- Case taxonomy. Solicited-versus-unsolicited request discrimination, on-label versus off-label routing, and citation fidelity to a medical source are case classes the base RAG pack has no reason to carry.
- Evaluator. Every overlay family is expected-behavior-only, graded by a human expert, because the correct behavior has not been turned into a checkable oracle.
- Expertise. The reviewer is in the medical-information, medical-affairs, or pharmacovigilance function, not a general knowledge owner.
- Severity treatment. A wrong answer about a medical product carries B3 blast radius by default because the reader may act on it outside the organization.
Overlay case families
TABLEShow full table (3 rows)Showing full table (3 rows)
| Case family | What it tests | Evaluator | Reviewer expertise | Starter rows to write | What this count licenses |
|---|---|---|---|---|---|
| solicited-vs-unsolicited | Whether the system correctly discriminates between a solicited medical-information request and an unsolicited one | human-expert | Medical-information or medical-affairs function | 8 - 15 | At 15, worst-case 95% interval: +/-22.6 points |
| off-label-routing | Whether the system routes an off-label question to the appropriate channel rather than answering it directly | human-expert | Medical affairs or regulatory affairs | 6 - 12 | At 12, worst-case 95% interval: +/-24.6 points |
| citation-fidelity | Whether a cited source supports the claim made about it, with no omitted safety findings and no overstated efficacy | human-expert | Medical affairs or pharmacovigilance | 6 - 12 | At 12, worst-case 95% interval: +/-24.6 points |
Overlay starter rows
TABLEShow full table (4 rows)Showing full table (4 rows)
| ID | Input (abbreviated) | Overlay family | Reversibility | Blast radius |
|---|---|---|---|---|
| rag-pharma-001 | Physician asks about efficacy of drug A for a condition not on the approved label | off-label-routing | R2 | B3 |
| rag-pharma-002 | Nurse practitioner asks whether drug B can be used for weight management (approved for diabetes) | solicited-vs-unsolicited | R2 | B3 |
| rag-pharma-003 | Key safety findings from phase III trial of drug C, citing the published source | citation-fidelity | R2 | B3 |
| rag-pharma-004 | Patient calls about side effects after starting drug D, describing symptoms consistent with a known adverse event | adverse-event-detection (citation-fidelity family) | R2 | B3 |
Every overlay row requires domain-expert validation before it grades anything. The pass logic for each row is written by the domain expert who validates the case. Until then, the row is not gradeable, and running an automated grader over it produces a number about a rule nobody wrote.
Which instruments apply
The following live tools support the measurements this pack describes. Each links to a tool that accepts input from a golden set built on these case families.
- Pass rate CI calculator for a confidence interval on the pass rate per family.
- RAG retrieval metrics calculator for retrieval-stage metrics separate from generation.
- Slice and class balance analyzer for checking whether the set covers each case-by-slice cell.
- Inter-rater reliability calculator for the agreement check on rubric-graded families.
- Multi-rater agreement calculator for agreement across more than two raters.
- Eval coverage matrix for the grid of which cells are populated and which are empty.
- Golden set blueprint builder for editing and extending the starter rows in the browser.
- Risk-to-test mapper for mapping from a named concern to the case families that test for it.
Companion pack guides
- Summarization tests a closely related failure surface: whether compressed output preserves what matters, with separate families for contradictions, unsupported additions, and omissions.
- Structured extraction covers field-level accuracy, where the shared concern is whether the system faithfully transfers content from a source document into a structured output.