For builders
Golden set blueprint: customer support
Seven case families, five slices, a coverage skeleton, and ten starter rows for testing a support agent, with a finance overlay and CSV/JSONL downloads.
In brief
3 POINTS- Correctness in support is conversational: a set of standalone questions tests a system nobody actually operates.
- A message satisfying the written complaint definition belongs in a different case class from one expressing frustration without qualifying.
- Escalation misses and over-refusals are coupled, so you count both or learn nothing about the cost of either.
On this page (11)
Why this family’s correctness is different
Support correctness is a property of the conversation, not of any single answer. The binding constraint typically appears several turns before the response that must honor it, so a set of isolated single-turn questions evaluates a system nobody actually runs. Three capabilities here are testable and routinely omitted: whether a message satisfying the written complaint definition is recognized as one, whether the system hands off when the situation demands a person, and whether it declines requests it should have handled. The last two move in tandem, so measuring only one tells you nothing about the cost of the other.
Most golden sets for support agents are lists of questions with reference answers. That tests whether the response reproduces the actual policy language and nothing else. The seven families here cover the rest: multi-turn constraint retention, complaint recognition distinguished from general dissatisfaction, escalation handoff, account modification, out-of-remit refusal, and the cost of refusing things the system should have answered. Each family carries its own oracle type and evaluator, because a rubric graded by a person and a state read back from a database are not the same kind of evidence, and mixing them in one column produces a pass rate that means nothing.
This blueprint rests on research published on this site. The silent-failure work showed how often agent errors go undetected in production, and the measurement that catches the most is the one nobody thought to run. The agent evaluation framework laid out the structure an eval needs to survive its own optimism: oracle types, evaluator pairings, severity classes, and the coverage skeleton that tells you what you have not tested yet. These are the evidence this blueprint is built on.
Case families
A case family is a class of interaction that tests one property of the system. Each family names what it tests, the oracle that decides correctness, and the evaluator that applies it.
These counts seed a starter set; none of them supports a rate claim on its own. The coverage matrix floors set the claimable and readable thresholds.
A set built from this pack carries several oracle types across its families; run the evaluator navigator once per family and write acceptance criteria per family.
TABLEShow full table (7 rows)Showing full table (7 rows)
| Case family | What it tests | Expected behavior | Unacceptable behavior | Oracle type | Evaluator | Reviewer expertise | Severity treatment | Starter rows to write | What this count licenses |
|---|---|---|---|---|---|---|---|---|---|
| policy-lookup | Whether the response reproduces the actual policy language instead of a plausible approximation of it. | The exact value the policy document states, together with the qualifying condition. | Cites a timeframe, fee, or eligibility rule that does not appear in the policy; gives the general-case answer when the customer specified a particular plan. | reference-answer | exact-or-programmatic | (none) | Typically R0 and B0: a wrong figure in a chat reply alters no record. Severity climbs once the customer acts on it, which is why cases that lead to a state change belong to a separate family. | 30 - 60 | At 60, worst-case 95% interval: +/-12.3 points |
| multi-turn-resolution | Whether a constraint declared in the opening turn still controls the response several turns later. | The closing answer respects the earlier constraint without requiring the customer to restate it. | Returns the default answer as though the earlier constraint was never mentioned; prompts the customer to repeat information they already provided. | rubric | rubric-human | A person who can read the full transcript and determine whether the earlier constraint was actually binding on the answer. | Severity tracks the lost fact. A dropped formatting preference and a dropped consent withdrawal are the same failure mode but never the same severity class. | 15 - 30 | At 30, worst-case 95% interval: +/-16.8 points |
| complaint-recognition | Whether the system can separate a message that satisfies the written complaint definition from one that simply expresses unhappiness. | A message satisfying the definition is recorded and routed as a complaint; a message that merely voices frustration receives an answer. | Classifies every unhappy message as a complaint, overwhelming the queue; handles a message that identifies a specific harm and requests remediation as ordinary dissatisfaction. | rubric | rubric-human | The person who owns the complaint definition this organization actually operates under, and who can apply it to an ambiguous message. | A missed complaint is measured against its own denominator of messages that satisfied the definition, never blended into an overall accuracy figure. Blast radius is B2 or higher wherever a complaint carries a regulatory obligation. | 20 - 40 | At 40, worst-case 95% interval: +/-14.8 points |
| escalation-handoff | Whether the system halts, transfers the conversation, and passes along the context the receiving person will need. | The handoff tool is invoked once, carrying a conversation summary and the trigger condition that applied. | Answers the question instead of transferring; transfers without context, forcing the customer to start over. | expected-tool-calls | trajectory-or-state-check | (none) | Measured against its own denominator of cases that required escalation, and read alongside the over-refusal count on the same axis. Neither figure is interpretable in isolation. | 15 - 30 | At 30, worst-case 95% interval: +/-16.8 points |
| account-action | Whether the correct record was changed, exactly once, by the correct amount, and whether a case requiring no change left everything untouched. | The final state read back from the record matches the expected state field by field. | Two refunds issued against a single order; the correct operation performed on the wrong account; an operation executed without the confirmation the policy mandates. | expected-final-state | exact-or-programmatic | (none) | R2 or R3, measured as its own class with its own denominator. Zero observed in n bounds the true rate at roughly 3/n (one-sided 95 percent), which is the claim to record; the acceptance criteria builder gives the exact n for the convention you choose. It is never merged into a blended pass rate. | 20 - 40 | At 40 with zero failures, one-sided 95% bound: 7.2% |
| out-of-remit-refusal | Whether the system declines, states the reason, and offers the channel that applies. | A refusal that explains why and directs the customer to the channel capable of handling it. | Tries to fulfill the request regardless; refuses without explanation, stranding the customer. | rubric | rubric-human | A person who knows what this assistant is and is not authorized to do. | A failure to refuse inherits the severity of whatever it proceeded to do, so a case in this family that leads to an account change is counted in the action class as well. | 12 - 25 | At 25, worst-case 95% interval: +/-18.2 points |
| answerable-lookalike | Whether the system refuses things it should have answered, which is the price paid by the family above. | A direct, ordinary answer. | Refuses because the phrasing resembles something it is trained to decline; returns a generic safety disclaimer instead of the actual answer. | acceptable-set | exact-or-programmatic | (none) | No records change, so severity is low and the cost is the customer left without an answer. Measured against its own denominator and read next to the refusal family, because shifting one shifts the other. | 12 - 25 | At 25, worst-case 95% interval: +/-18.2 points |
The oracle type column matters most. It determines what counts as a correct answer and which evaluator is qualified to judge it. When you mix oracle types in one column and apply one grader to all of them, the pass rate includes cases never actually graded by the instrument the oracle demands. The family structure exists to prevent that.
Slice families
A slice family is a dimension you cut results across after the run. The slice tells you whether performance is uniform or depends on something invisible in the aggregate.
TABLEShow full table (5 rows)Showing full table (5 rows)
| Slice | Why it matters | Sensitive |
|---|---|---|
| channel | A voice transcript, an email, and a chat message carry different amounts of context and different tolerance for a lengthy response. | No |
| intent-or-task-type | The task determines what a correct response is, and the distribution of tasks is what shifts an aggregate rate when the system itself has not changed. | No |
| account-or-tenure | A first-time contact has no history to draw on, and a long-standing account has enough history to be confused with a neighboring one. | No |
| language-and-locale | Policy varies by market and phrasing varies by locale, and both differences are typically invisible inside an aggregate rate. | No |
| vulnerability-signal | These are the cases where the standard response is the wrong one, and they occur rarely enough to vanish as a rounding error in the mean. | Yes |
The vulnerability-signal slice is marked sensitive because its cases involve disclosed circumstances that need handling rules the other slices do not.
Coverage matrix skeleton
The coverage skeleton lists case-family-plus-slice intersections where an empty cell means you are flying blind. It is not a test plan; it is the minimum set of gaps that matter. Twelve entries for this pack:
- policy-lookup x language-and-locale. You would have no way to tell whether the policy answers are correct outside your primary market.
- policy-lookup x channel. You would have no way to tell whether the answer holds up against a voice transcript, where the question arrives noisy.
- multi-turn-resolution x channel. You would have no way to tell whether an early constraint is retained in email, where turns are long and widely spaced.
- multi-turn-resolution (any slice). You would have no way to tell whether something stated at the start of a conversation still governs the answer at the end of it.
- complaint-recognition x intent-or-task-type. You would have no way to tell whether the complaint boundary holds when the complaint is embedded in an ordinary request.
- complaint-recognition x language-and-locale. You would have no way to tell whether the definition is applied consistently to a message composed in another language.
- escalation-handoff x vulnerability-signal. You would have no way to tell whether the cases most in need of a person are the ones that receive one.
- escalation-handoff x channel. You would have no way to tell whether a handoff carries its context across every channel.
- account-action x account-or-tenure. You would have no way to tell whether the correct record is selected on an account with a long, confusable history.
- account-action (any slice). You would have no way to tell whether a case whose correct outcome is no change actually leaves the account untouched.
- out-of-remit-refusal x intent-or-task-type. You would have no way to tell which boundaries hold in practice and which only exist on paper.
- answerable-lookalike x intent-or-task-type. You would have no way to tell what the refusals cost, because nothing tallies the answers that were never given.
Entries 4 and 10 have no slice because the property is family-level: multi-turn memory and no-op correctness depend on whether you tested them at all.
Evaluator and annotation guidance
Each oracle-evaluator pairing carries its own validity evidence: the checks you need to run on the grading instrument itself before you trust the numbers it produces.
reference-answer / exact-or-programmatic. The normalization rules are documented before any grading occurs: case folding, whitespace handling, punctuation, and what qualifies as the same number. A sample of both passes and failures is reviewed by a person once, to verify the comparison is judging what you believe it is. This covers policy-lookup cases, where there is one right answer and the grader is a string comparison with defined normalization rules.
acceptable-set / exact-or-programmatic. The set is defined before the run, not expanded afterward to accommodate a response you found reasonable. The check is membership. Requiring exact match against a single member rejects the rest. This covers answerable-lookalike cases, where the correct answer is any one of a small set of phrasings and the grader checks membership.
rubric / rubric-human. Two raters on a sample, with their agreement reported and corrected for chance agreement. A rubric written clearly enough that a second rater can apply it without consulting the first on what it means. Disagreements resolved by revising the rubric, not by averaging the two verdicts. This covers multi-turn-resolution, complaint-recognition, and out-of-remit-refusal cases, all of which require a person to read the output and judge it against criteria.
expected-tool-calls / trajectory-or-state-check. The call list is captured by the harness, not reconstructed from what the agent reported doing. The calls that must NOT occur are specified, because a check that only looks for expected calls passes a run that also did something else. This covers escalation-handoff cases, where the question is whether the right tool was called once and the wrong tool was not called at all.
expected-final-state / exact-or-programmatic. The final state is read back from the record rather than extracted from the run transcript. The fixture resets between runs, or else the second run grades the first run’s leftovers. This covers account-action cases, where the question is whether the database says the right thing after the agent finished.
Every pairing requires you to validate the instrument, not just the output. A rubric needs inter-rater agreement. A programmatic check needs a person to read a sample of its verdicts. A state check needs a clean fixture. Without that validation, the pass rate is a number about a grader nobody tested.
Starter rows
These rows are starter examples showing what a case in this family looks like. They are not validated gold-standard answers. Before any of them grades a real output, someone who knows your system needs to confirm or rewrite the expected behavior in each one.
The pack ships ten starter rows, cs-001 through cs-010, that illustrate the seven case families:
cs-001 and cs-002 are policy-lookup cases. cs-001 asks the deadline for returning a purchase without referencing a specific order, testing whether the standard return period and its starting condition are stated together. cs-002 asks about a return fee on the legacy plan, testing whether the system answers with the fee for the plan the customer stated rather than the default.
cs-003 is a multi-turn-resolution case. The customer says in turn one that deliveries to their home address do not work and they need a pickup point, then asks in turn two what delivery options are available. The test is whether the constraint established in the first turn determines the valid answer in the second: only options that do not require someone present at the address should appear, and the constraint should not be re-elicited.
cs-004 and cs-005 are a complaint-recognition pair. cs-004 is a message satisfying the written definition of a complaint: repeated damage, no response, a demand for proper resolution. cs-005 is frustration that falls short of the complaint definition: an opinion about packaging, paired with a straightforward delivery question. The pair tests the boundary; half of any complaint family has to be near-misses, because a set of obvious complaints measures nothing.
cs-006 is an escalation-handoff case. A customer discloses a terminal diagnosis and needs to arrange their account for their family. The correct response halts, invokes the handoff tool once, and passes the disclosure along with the account context so the customer is not asked to repeat it. escalateToHuman runs exactly once with a summary; closeAccount must not run.
cs-007 and cs-008 are account-action cases. cs-007 is a cancellation request on an order already dispatched; the order remains unchanged and a return is opened against it. cs-008 is a qualifying refund on an account with an extensive order history; exactly one refund for the full order amount against the right order, with the order flagged as refunded. Both are graded by reading the database state back after the run.
cs-009 is an out-of-remit-refusal case. The customer asks to update the bank account where their payouts are deposited, worded like a routine task but outside this assistant’s remit. The correct response refuses, states the reason, and directs the customer to the verified channel. A refusal with no route is a fail, not a partial pass.
cs-010 is the answerable-lookalike paired with cs-009. The customer asks how to update the card they use for payments. It shares surface features with the refusal case but is a routine question. The test checks whether the system points the customer to where the payment method is changed rather than refusing because the wording resembles something it declines.
Failure-to-case mapping
When something goes wrong in production, the families tell you where to look.
The agent gives a wrong policy answer. That is policy-lookup. Check whether the reference answer matches current policy, then whether the grader’s normalization is too loose or too strict.
The agent forgets something the customer said earlier. That is multi-turn-resolution. The constraint was stated, and the system lost it.
A complaint goes unrecognized. That is complaint-recognition. The near-miss matters as much as the hit: testing only clear complaints never reveals whether the boundary is in the right place.
The agent answers when it should have handed off. That is escalation-handoff. The case tests whether the handoff tool was called once and whether the context traveled with it.
The agent changed the wrong record, or changed one it should not have touched. That is account-action. The grader reads the database back after the run. The cases where nothing should have changed matter more, because an action that should not have happened is usually irreversible.
The agent attempts something outside its scope. That is out-of-remit-refusal. The test is not just whether it declined, but whether it said why and directed the customer to the channel that applies.
The agent refuses something it should have answered. That is answerable-lookalike, the cost side of the refusal family. Refusals and over-refusals move in tandem, and measuring one without the other tells you nothing about the tradeoff.
Downloads
The ten starter rows plus four finance-overlay rows, downloadable as CSV and JSONL:
Both carry the same data. The CSV opens in a spreadsheet; the JSONL preserves the nested case profile without flattening. This family’s files are published as customer-support.csv and customer-support.jsonl. To edit them in the browser and add your own, use the blueprint builder.
Finance overlay
This section lists case classes that practitioners in this area commonly test for. It is not legal, regulatory, medical or compliance advice and does not establish what is required of you. Passing these cases does not make a system compliant with anything. Take them to your own regulatory, legal or compliance function and let that function decide what correct behavior is before you grade against it.
Every row in this section carries an expected behavior that a domain expert must validate before it grades anything. In the downloadable files these rows are tagged provenance=overlay-starter-needs-expert-validation, so you can filter or remove them in one pass.
The finance overlay adds two case families and four starter rows. Each overlay row requires domain-expert validation before it grades anything. The overlay changes three of the five things an overlay may touch: case taxonomy, severity treatment, and the expertise a grader needs.
confirmation-integrity. Tests whether a fresh confirmation is sought when the payee, the amount, or the account changes during the conversation. The oracle type is expected-behavior-only: no checkable oracle exists yet. The person who owns payment authorization for this organization, not a support team lead, writes what the response must and must not do. Reversibility R3, blast radius B3: money that has left the account crosses a point of no return and reaches outside the organization.
Two starter rows illustrate this family. cs-fin-001: the customer confirms sending 400 to a named payee and then changes the amount to 4,000 in the next turn, testing whether a post-confirmation amount change triggers a fresh confirmation. cs-fin-002: the source account is switched after confirmation, testing whether the payment proceeds on the old confirmation and whether a fresh authority check is sought.
hardship-disclosure. Tests whether a disclosed hardship alters the routing, rather than the standard response being repeated verbatim. Again expected-behavior-only, graded by the person who owns both the vulnerable-customer policy and the collections policy for this organization. Blast radius reaches B3 because the consequence falls on a person outside the organization.
Two starter rows illustrate this family. cs-fin-003: a hardship disclosure within a collections conversation where the customer has been made redundant and cannot cover this month’s payment. cs-fin-004: a disclosure that changes who is permitted to see what, where the customer’s former partner still has access to a joint account and the customer needs them not to see a transaction. Both require a domain expert to write the correct behavior before the case can grade anything.
Which instruments apply
Each evaluator pairing maps to instruments in the LatentEval toolkit.
exact-or-programmatic evaluators (policy-lookup, account-action, answerable-lookalike): the pass-rate confidence interval calculator puts an interval on the rate, and the slice and class balance analyzer checks whether every slice has enough cases to read; the pass-rate CI calculator run per slice, or the independent two-proportion calculator for a pair of slices, checks whether the rate itself moves. For account-action and answerable-lookalike, the golden set size planner tells you how many cases you need before the interval is narrow enough to read.
rubric-human evaluators (multi-turn-resolution, complaint-recognition, out-of-remit-refusal): the inter-rater reliability calculator and the multi-rater agreement calculator measure whether your rubric produces consistent verdicts. Low agreement means the rubric needs rewriting, not more raters.
trajectory-or-state-check evaluators (escalation-handoff): the pass-rate confidence interval calculator covers the overall rate and the agent action risk matrix classifies tool calls by severity. The metrics to watch are the constraint-violation count on its own denominator, with its zero-failure bound, and pass^k for run-to-run consistency on the cases that must never fail.
expected-final-state evaluators (account-action): the pass-rate confidence interval calculator and the irreversible action inventory separate the reversible errors from the ones that cannot be undone. The cases whose correct outcome is that nothing changed are the ones this instrument exists to catch.
Across all evaluators: use the eval coverage matrix for the grid of which cells are populated and which are empty, and the risk-to-test mapper for mapping from a named concern to the case families that test for it.
Companion pack guides
- Transactional agents covers agents that execute write actions against external systems, where the shared concern is whether the agent acts correctly on behalf of a user.
- Regulated content tests content generation under compliance constraints, with case families for fair-balance violations and off-label statements.