For builders
Golden set blueprint: research agents
Seven case families for research agents: citation existence, citation support, evidence strength, source coverage, and silent fabrication, each graded per claim against a named source.
In brief
3 POINTS- You check support per claim against a named source and aggregate afterward, never as a single label on the whole response.
- Citation existence, attribution fidelity, and evidence-strength accuracy are three separate case families with separate evaluators.
- Silent success is the headline risk: a confident, well-sourced report that nothing in your test set catches.
On this page (10)
Why this family’s correctness is different
Correctness here is per-claim, not per-response. A research agent that produces a confident, well-sourced report can still be wrong on half of its citations, and a response-level pass rate will not tell you, because the response reads fine. The test is whether every claim traces to a named source, whether the source actually says what the claim says it does, and whether the strength of the claim is proportional to what the source supports. Those are three separate axes, not one. An agent that cites a real paper, quotes it faithfully, and overstates its conclusion fails on the third without failing on the first two, which is why this blueprint treats them as separate case families with separate evaluators.
The headline risk is silent success. The output is fluent, the citations look right, the hedging reads as appropriately cautious, and nothing in the response signals that a claim was fabricated, a source was misread, or a preliminary finding was reported as settled science. You cannot catch this at the response level. You catch it by checking each claim against its named source, independently, and aggregating afterward.
This blueprint rests on primary research published on this site. The silent-failure work showed that the most dangerous agent errors are the ones that look like correct outputs to every downstream consumer, and that a confident, well-formatted response is no evidence that the underlying claims hold. The broader agent evaluation framework established the structure an eval needs to survive: oracle types, evaluator pairings, severity classes, and the coverage skeleton that tells you what you have not tested yet. These are the evidence this blueprint is built on.
Case families
A case family is a class of test case that shares an oracle type and an evaluator. The seven families below cover the failure modes that matter most for a research agent. Each family gets its own denominator. You do not fold “does this citation point to a real document” into the same pass rate as “does the agent inflate the certainty of a preliminary finding,” because those are different instruments measuring different properties.
These counts seed a starter set; none of them supports a rate claim on its own. The coverage matrix floors set the claimable and readable thresholds.
A set built from this pack carries several oracle types across its families; run the evaluator navigator once per family and write acceptance criteria per family.
TABLEShow full table (7 rows)Showing full table (7 rows)
| Case family | What it tests | Expected behavior | Unacceptable behavior | Oracle type | Evaluator | Reviewer expertise | Severity treatment | Starter rows to write | What this count licenses |
|---|---|---|---|---|---|---|---|---|---|
| citation-existence | Whether every source the agent names is a real, retrievable document rather than a plausible-sounding fabrication. | Every citation resolves to a document that can be independently retrieved. No source is invented. | A citation points to a document that does not exist; a citation names a real journal and a fabricated article; a URL leads to a page unrelated to the claim. | reference-answer | exact-or-programmatic | (none) | R2: a fabricated citation that a reader acts on cannot be undone by the reader, because they have no way to discover it is fabricated without checking. Count fabricated citations as their own class with their own denominator, never folded into a general accuracy rate. | 10 - 20 | At 20, worst-case 95% interval: +/-20.1 points |
| citation-support | Whether the claim the agent attributes to a source is actually present in that source, as opposed to a paraphrase that drifts from the original meaning. | Each claim attributed to a source can be traced to a specific passage, and the passage supports the claim as stated. | The source discusses the topic but does not make the specific claim attributed to it; the claim inverts the source (e.g. the source says X is unlikely, the claim says X is likely); the source is about a different population or context than the claim implies. | rubric | rubric-human | Someone who can read the source material in its original form and judge whether a paraphrase preserves or distorts meaning. | R2 when the distortion changes direction (a source that hedges, cited as definitive), R1 when the distortion is a matter of emphasis. Count direction-changing distortions as their own class. | 15 - 30 | At 30, worst-case 95% interval: +/-16.8 points |
| citation-strength | Whether the agent inflates or deflates the certainty of a finding relative to what the source actually establishes. A preliminary finding reported as settled science, or a strong consensus reported as tentative, both fail here. | The hedging, qualifiers, and certainty language in the output match or conservatively follow the hedging in the source. Preliminary findings are described as preliminary. Consensus findings are described as established. | A single small study cited as though it established a consensus; a well-replicated finding dismissed as inconclusive; correlation language in the source rendered as causal language in the output. | rubric | rubric-human | Someone who can distinguish between levels of evidence strength in the relevant domain and can read hedging language in the source. | R1: the report can be corrected. B2 when a reader would act differently on the inflated claim than on the actual finding, B1 when the inflation is a matter of style rather than decision-relevance. Count inflation and deflation as separate classes. | 12 - 25 | At 25, worst-case 95% interval: +/-18.2 points |
| source-coverage | Whether the agent retrieved the sources that matter for the question, rather than settling for sources that are easy to find and partially relevant. | The retrieved set includes the key sources a competent researcher would find, identified before the run. Sources may be omitted from the final output as long as the agent considered and triaged them. | A key source was available and the agent did not retrieve it; the agent retrieved only sources that support one side of a contested question; the agent stopped searching after the first plausible result. | acceptable-set | exact-or-programmatic | (none) | R1: the report can be rerun with broader retrieval. B2 when the missing source would have reversed the conclusion, B1 when it would have added nuance. Count missing-key-source as its own class. | 10 - 20 | At 20, worst-case 95% interval: +/-20.1 points |
| silent-fabrication | Whether the agent generates claims that read as sourced but are not attributable to any source in the retrieved set or the agent conversation. This is the headline risk: silent success, where the output is confident, fluent, and wrong. | Every factual claim in the output is either attributed to a named source or flagged as the agent’s own inference. No unsourced claim is stated with the same confidence as a sourced one. | A specific statistic appears with no source and no qualifier; a factual claim is presented as sourced when no retrieved document contains it; the agent interpolates between two sources and presents the interpolation as a finding. | rubric | llm-judge | (none) | R1: the report can be corrected once the fabrication is discovered. B3 when the fabricated claim is actionable (a statistic a decision-maker would use), B2 for a fabricated contextual detail. Count fabrications as their own class, never folded into a general quality rate. | 12 - 25 | At 25, worst-case 95% interval: +/-18.2 points |
| contradictory-sources | Whether the agent handles genuine disagreement between sources by reporting the disagreement rather than picking a side silently or averaging the positions into a false consensus. | The output names the disagreement, attributes each position to its source, and does not assert a resolution the sources do not support. | The agent picks one source and ignores the other without explanation; the agent merges two contradictory positions into a blended claim that neither source makes; the agent notes the disagreement but resolves it by citing the more recent source without explaining why recency matters here. | rubric | rubric-human | Someone who can read both sources and judge whether the agent faithfully represented the nature and scope of the disagreement. | R1: the report can be corrected. B2 when the suppressed disagreement would change a decision, B1 when the disagreement is real but immaterial to the question asked. Count cases where a disagreement was hidden as their own class. | 8 - 15 | At 15, worst-case 95% interval: +/-22.6 points |
| synthesis-fidelity | Whether the agent synthesizes findings from several sources without distorting any individual source in the process. A synthesis that reads smoothly but misrepresents one contributor fails here. | The synthesized statement is traceable back to each contributing source, and no contributing source is misrepresented by the synthesis. | A synthesized claim overgeneralizes one source to align it with the others; a dissenting source is included in the bibliography but its dissent is absent from the synthesis; the synthesis implies agreement between sources that addressed different questions. | rubric | rubric-human | Someone who can read each source individually and then judge whether the combined statement faithfully represents all of them. | R1: the report can be corrected. B2 when the distorted synthesis would lead to a different action than a faithful one, B1 when the distortion is a simplification that preserves the decision-relevant content. Count synthesis distortions as their own class. | 8 - 15 | At 15, worst-case 95% interval: +/-22.6 points |
The oracle type column determines what counts as a correct answer. Citation-existence and source-coverage can have checkable oracles once you build the reference set; the starter rows ship as expected-behavior-only because no reference set exists yet, and they admit only a human expert until you do. The other five are rubric-graded, because the judgment is about fidelity, strength, and faithful representation, none of which reduce to a string match.
Slice families
A slice family is a dimension you cut your results across after the run. Every case belongs to one or more slices, and the slice tells you whether performance is uniform or whether it depends on something you did not notice in the aggregate.
TABLEShow full table (5 rows)Showing full table (5 rows)
| Slice | Why it matters | Sensitive |
|---|---|---|
| content-or-document-type | A journal article, a policy document, a dataset, and a web page carry different conventions for hedging, citation, and evidence strength. An agent that handles one well may misread the norms of another. | No |
| input-length | A single-document question and a multi-document synthesis exercise different failure paths. Fabrication and omission both rise with the volume of material the agent is expected to cover. | No |
| intent-or-task-type | A fact-check, a literature review, and a source comparison require different citation behaviors. A fact-check needs one authoritative source; a literature review needs breadth and balance. | No |
| upstream-source | Academic databases, web search results, and internal document stores differ in reliability, recency, and the likelihood of returning fabrication-adjacent near-misses. An agent that works well on curated sources may fabricate when the retrieval pipeline returns noise. | No |
| time-or-recency | A question about current guidance and a question about historical findings put different pressure on the retrieval pipeline. Stale-source failures concentrate in recency-sensitive queries. | No |
None of these slices are marked sensitive. The data in a research-agent golden set is about publications and sources, not about individuals or protected characteristics.
Coverage matrix skeleton
The coverage skeleton is a list of case-family-plus-slice intersections that tells you what you have not tested. Each entry names a case family, optionally pairs it with a slice, and states what you would not know if that cell were empty. It is not a test plan. It is the minimum set of intersections where an empty cell means you are flying blind on something that matters.
The twelve entries for this pack:
- citation-existence x upstream-source. You would not know whether fabrication rates differ by source pipeline, which is where the mechanism lives.
- citation-existence x content-or-document-type. You would not know whether the agent fabricates more in one document type than another.
- citation-support x intent-or-task-type. You would not know whether attribution fidelity holds across fact-checks, reviews, and comparisons.
- citation-strength x content-or-document-type. You would not know whether the agent inflates certainty more with one kind of source than another.
- source-coverage x input-length. You would not know whether the agent stops searching earlier when there are many sources to find.
- source-coverage (any slice). You would not know whether the agent retrieves the sources that matter, or just the ones that are easy to find.
- silent-fabrication x upstream-source. You would not know whether fabrication concentrates in cases where the retrieval pipeline returns thin results.
- silent-fabrication x time-or-recency. You would not know whether the agent fabricates more when asked about recent developments where its training data is thin.
- contradictory-sources (any slice). You would not know whether the agent surfaces disagreement or silently picks a side.
- contradictory-sources x content-or-document-type. You would not know whether the agent handles disagreement differently in clinical literature than in policy documents.
- synthesis-fidelity x input-length. You would not know whether synthesis quality degrades as the number of sources grows.
- citation-strength x time-or-recency. You would not know whether the agent inflates the certainty of recent, preliminary findings more than it inflates established ones.
Entries 6 and 9 have no slice because the property they test is family-level: source retrieval quality and disagreement handling do not depend on a dimension, they depend on whether you tested them at all.
Evaluator and annotation guidance
Each oracle-evaluator pairing carries its own validity evidence: the checks you need to run on the grading instrument itself before you trust the numbers it produces.
reference-answer / exact-or-programmatic. Used for citation-existence and some source-coverage checks. Establish the reference set before the run, not from the agent output. Run existence checks against an independent retrieval path, not the same pipeline the agent used. The check distinguishes a real source with a wrong detail (wrong year, wrong author order) from a fabricated source. A paper with the right title but the wrong journal is not the same finding as a paper that does not exist at all.
acceptable-set / exact-or-programmatic. Used for source-coverage cases where the question is whether the agent retrieved the key sources a competent researcher would find. The acceptable set was curated by someone who knows the field, not generated by the system under test. The set includes sources the agent might plausibly miss, not only the obvious ones. Membership is checked by content identity, not by exact title match, because titles vary across databases.
rubric / rubric-human. Used for citation-support, citation-strength, contradictory-sources, and synthesis-fidelity. The rubric criteria were written against the source material before the agent output was seen. At least two raters scored a sample of cases and their agreement was measured before the full set was graded. The rubric distinguishes citation existence, citation support, and citation strength as separate dimensions, and the rater applies only the dimension this family tests. Disagreements are resolved by rewriting the rubric, not by averaging the two verdicts.
rubric / llm-judge. Used for silent-fabrication, where the volume of claims makes human grading at scale expensive. The judge was validated against human labels on items from this domain, not a generic benchmark. The validation set included cases where the source exists but does not support the claim, which is the hard instance for judges. Agreement between the judge and human raters was measured on both positive and negative cases. An LLM judge with no validation against human verdicts on research claims is a second opinion from the same kind of system you are evaluating.
Starter rows
These rows are starter examples showing what a case in this family looks like. They are not validated gold-standard answers. Before any of them grades a real output, someone who knows your system needs to confirm or rewrite the expected behavior in each one.
The pack ships ten starter rows (ra-001 through ra-010) that illustrate the case families and slices above.
ra-001 asks what the current evidence says about intermittent fasting and cardiovascular risk in adults over 50. This is a broad literature review question with multiple RCTs and meta-analyses available, some contradicting each other. The test checks whether at least three primary sources are cited, each of which exists and supports the claim attributed to it, and whether contradictory findings are surfaced rather than merged. Sliced on journal-article and literature-review.
ra-002 asks for the three most-cited papers on transformer attention mechanisms published before 2020. A factual retrieval task where the answer is verifiable against citation databases. Each paper title must resolve to a real publication, publication dates must be before 2020, and citation counts must be independently verifiable. The agent must not invent titles or counts. Sliced on journal-article and fact-check.
ra-003 asks what the IPCC AR6 Working Group I report says about the likelihood of exceeding 1.5 degrees C by 2040. A single-source question where the risk is that the agent replaces the IPCC’s calibrated likelihood language with unqualified certainty. The test checks whether the hedging matches the IPCC calibrated scale and whether no claim exceeds the strength of the source language. Sliced on policy-document and fact-check.
ra-004 asks whether studies show that code review by AI tools reduces defect density. A question in a fast-moving field where few rigorous studies exist. The test checks whether the agent reports the thinness of the evidence base rather than filling gaps with plausible-sounding claims, and whether every named study is retrievable with its limitations mentioned. Sliced on journal-article and compare.
ra-005 asks the agent to compare two studies on sleep duration and memory consolidation that reach different conclusions. The test is whether both studies are described faithfully and the disagreement is named and attributed to methodological differences rather than resolved by picking a side. Sliced on journal-article and compare.
ra-006 asks what the recent literature says about large language models as judges in NLP evaluation. A broad question where the risk is that the agent cites only sources supporting one view. The test checks membership against a pre-established acceptable set of key papers and verifies each attribution. Sliced on web-page and literature-review.
ra-007 asks whether there is evidence that microplastics accumulate in human lung tissue. A factual question where a small number of primary studies exist. The test checks that each cited study is retrievable, is about human lung tissue specifically (not animal tissue), and is primary research rather than secondary commentary. Sliced on journal-article and fact-check.
ra-008 asks the agent to synthesize findings from three provided papers on remote work productivity. Three papers with different methodologies and partially conflicting results. The test checks that each paper is individually accurately represented, that methodological differences are named, and that the overall conclusion does not exceed what the three sources jointly support. Sliced on dataset and compare.
ra-009 asks for the current WHO guidance on COVID-19 booster vaccination schedules as of 2024. A question where the answer has changed over time. The test checks that the cited document is the current one, not an earlier version, and that superseded guidance is not presented as current. Sliced on policy-document and fact-check.
ra-010 asks for a summary of legal precedent on AI-generated content and copyright in the United States. A fast-moving legal area where rulings are recent. The test checks that each named case is real, the court and year are correct, the holding is accurately stated, and pending cases are distinguished from decided ones. Sliced on web-page and literature-review.
Failure-to-case mapping
When a research agent fails in production, the failure maps to one of the case families above. Knowing the mapping tells you which family needs more cases.
The agent cites a paper that does not exist. A plausible-sounding title in a real journal, with fictional authors or a fabricated DOI. Maps to citation-existence. Check whether the reference set in your golden set covers the document types where fabrication concentrates.
The agent says the study found X when it found Y. The source is real and the claim is attributed to it, but the paraphrase drifts from the original meaning, or the claim inverts the source’s direction. Maps to citation-support. The distinction matters: the agent misread what it found rather than fabricating content.
The agent states a definitive conclusion from a weak pilot study. A single small study cited as though it established a consensus, or correlation language in the source rendered as causal language in the output. Maps to citation-strength. This is the subtlest failure: the source is real, the claim is present in it, and the inflation happens in the hedging, not in the content.
The agent missed the key paper on the topic. The retrieved set is missing a source that a competent researcher would have found, and the missing source would have changed the conclusion. Maps to source-coverage. The severity follows what the missing source would have changed: a reversal is worse than added nuance.
The agent states a specific statistic with no source and no qualifier. A factual claim presented as sourced when no retrieved document contains it, or a number interpolated between two sources and presented as a finding. Maps to silent-fabrication. This is the headline risk: the output is confident, the claim is unsupported, and nothing in the response signals the gap.
The agent picks one source and ignores the contradicting one. Two sources disagree, and the agent merges them into a false consensus or silently drops the dissenting view. Maps to contradictory-sources. The test checks whether the agent reported the disagreement.
The synthesis misrepresents a contributing source to make the narrative smoother. A dissenting source is included in the bibliography but its dissent is absent from the synthesis, or one source is overgeneralized to align it with the others. Maps to synthesis-fidelity. The synthesis reads well, which is exactly why this family exists.
Downloads
The ten starter rows are downloadable as CSV and JSONL:
Both formats carry the same data. The CSV is easier to open in a spreadsheet; the JSONL preserves the nested case profile without flattening. This family’s files are published as research-agent.csv and research-agent.jsonl. To edit them in the browser and add your own, use the blueprint builder.
Which instruments apply
Each evaluator-guidance pairing above maps to specific tools and metric families in the LatentEval toolkit.
exact-or-programmatic evaluators (citation-existence, source-coverage): use the pass-rate confidence interval calculator to put an interval on the per-family rate, and the RAG retrieval metrics calculator to measure whether the retrieval pipeline surfaces the right sources. The primary metric families are pass-rate (did the citation check pass), per-claim-support (the per-claim breakdown rather than the per-response aggregate), and per-slice-rate (whether the rate holds across upstream-source and content-type slices).
rubric-human evaluators (citation-support, citation-strength, contradictory-sources, synthesis-fidelity): use the inter-rater reliability calculator and the judge agreement tracker to measure whether your rubric produces consistent verdicts across raters. If agreement is low, the rubric needs rewriting, not more raters. The primary metric families are per-claim-support (each claim graded individually), agreement (inter-rater on the rubric), and pass-rate (the family-level aggregate once the rubric is stable).
llm-judge evaluators (silent-fabrication): use the judge validation report builder to produce the validation evidence and the judge calibration calculator to check whether the judge’s confidence tracks its accuracy. The primary metric families are agreement (judge versus human on positive and negative cases), calibration (whether the judge knows when it is uncertain), and per-claim-support (the claim-level verdict, not the response-level one). Validate the judge on research claims from your domain before trusting it at scale.
Across all evaluators: use the eval coverage matrix for the grid of which cells are populated and which are empty, and the risk-to-test mapper for mapping from a named concern to the case families that test for it.
Companion pack guides
- RAG and knowledge assistants shares the retrieval-fidelity concern but tests it at the response level rather than the research-synthesis level.
- Summarization covers compression fidelity, where the failure families overlap with research synthesis: unsupported additions and salient omissions.