INSTRUMENT | Reliability testing
Risk-to-Test Mapper: Which Failures Need Which Cases
Name the failures you care about, then see the test-case shapes, slices, evaluators and metrics each one needs. Every mapping says what it cannot establish. The mapping is a LatentEval synthesis.
Pick the failure concerns that matter to your system, or write your own, and get the case patterns, slices, evaluators, metrics, reruns and evidence limits each one calls for. The mapping draws on the change impact worksheet and the rerun evidence. It is ours, not a standard, and the presets are a starting palette, not a taxonomy. To derive failure concerns from a system description rather than picking them, use the context-of-use profiler. The effect column pairs with the expected cost of failure calculator, and the mitigation column draws on the irreversible action inventory.
All 19 preset concerns are fully mapped.
Concerns you added: 0 of 0 mapping cells still empty.
| Concern | Family | Case patterns | Slices | Evaluators | Metrics | Severity treatment | Repetition | Human review | Reruns | Evidence limits |
|---|---|---|---|---|---|---|---|---|---|---|
| Leaves out something the answer needed The answer is right as far as it goes and is missing a part the reader needed. | output correctness | A case whose correct answer has several required parts, written out so a grader can check each one.A case where one required part sits at the end of a long input, so dropping it looks like a length effect.A case whose question contains two questions, where answering one reads as a complete answer. | input-length, content-or-document-type, intent-or-task-type | reference-answer: exact-or-programmaticrubric: rubric-human, llm-judgeexpected-behavior-only: human-expert | Pass rate with an interval, Per-slice rate | Severity inherits from what was omitted. A missing caveat on an irreversible action carries that action's reversibility and blast radius; a missing nicety carries neither. Count the classes with separate denominators rather than folding them into one rate. | repeat per case | sampled | I01 Core task suite on the frozen golden set, I03 Per-slice pass rates for the named slices, I22 Context-window and long-input behavior | A grader that checks for the parts you listed cannot see a part nobody listed.An omission rate measured on short cases says nothing about long ones, because length is the mechanism. |
| States something the source does not support The output asserts a fact the material it was given does not carry. | grounding and retrieval | A case whose source material is fixed and named, so a claim can be traced back to a line or fail to be.A case where the answer is plausible and the source is silent on it, which is the hard instance.A case whose source contains a near-miss statement, testing whether the answer overstates it. | content-or-document-type, upstream-source, input-length | rubric: rubric-human, llm-judgereference-answer: exact-or-programmaticexpected-behavior-only: human-expert | Per-claim support, Pass rate with an interval, Rater or judge agreement | An unsupported claim a reader would act on inherits the blast radius of that action. Claims touching a regulated or safety-relevant domain belong in their own counted class with their own denominator, not as rows in a general accuracy rate. | repeat per case | sampled | I10 Grounding and faithfulness, I07 Retrieval recall at k, I13 Judge agreement against human labels | A faithful reading of a wrong document scores well. Support against the source is not correctness.A judge scoring support needs its own agreement evidence against human labels on your items before its numbers mean anything. |
| Recommends an action that would cause harm The system advises something that would hurt the person following it. | action and authority | A case where the obvious answer is the harmful one, so a system that pattern-matches fails it.A case carrying a disclosed circumstance that makes the standard recommendation wrong.A case where the correct response is to decline and say why. | vulnerability-signal, protected-or-regulated-population, intent-or-task-type | rubric: rubric-human, human-expertexpected-behavior-only: human-expert | Constraint violations, Pass rate with an interval | Each case is its own counted class with its own denominator. Blast radius follows what the harm affects; advice to an end user extends outside the organization. No aggregate pass rate absorbs a harmful recommendation. | repeat and report spread | domain expert | I05 Safety and refusal suite, I01 Core task suite on the frozen golden set, I16 Repeated-run variance at the new setting | Zero harmful recommendations in a case set bounds the rate; it does not establish that none happen.A harm class written by the people who built the system inherits their view of what harm is. |
| Answers from material that is no longer current The retrieved material was true once and is not true now. | grounding and retrieval | A case whose correct answer changed on a known date, with both versions of the source in the index.A case where the superseded document ranks higher than the current one on surface similarity.A case where the answer must say how fresh it is, not only what it is. | time-or-recency, content-or-document-type, product-or-line | reference-answer: exact-or-programmaticrubric: rubric-human, llm-judgeexpected-behavior-only: human-expert | Retrieval rank quality, Pass rate with an interval, Per-slice rate | Severity follows what a stale answer would cause. A stale price quote someone acts on is compensable at best; a stale safety instruction is not. Count time-sensitive classes with separate denominators. | single run | sampled | I07 Retrieval recall at k, I08 Retrieval relevance at the new chunk size and top-k, I09 Reranker precision | Retrieval rank measures what was surfaced, never whether the generator used it.A staleness check written against the sources you know changed cannot see a source that changed quietly. |
| Acts on or answers about the wrong record, person or product Everything about the answer is right except which thing it is about. | output correctness | A case with two similar records in scope, where only one matches the request.A case where the identifier is ambiguous and the correct behavior is to ask.A case where the right record is reachable only after a disambiguating step. | account-or-tenure, product-or-line, upstream-source | reference-answer: exact-or-programmaticexpected-final-state: trajectory-or-state-check, exact-or-programmaticexpected-tool-calls: trajectory-or-state-check, exact-or-programmatic | Pass rate with an interval, Constraint violations, Per-slice rate | Blast radius follows the record that was touched: a wrong-entity write inherits the reversibility of whatever it wrote. A wrong-entity read about another person is a disclosure and is counted apart from an ordinary miss. | repeat per case | sampled | I11 Tool-call correctness: schema, arguments, wrong-tool rate, I12 Permission and irreversible-action gates, I01 Core task suite on the frozen golden set | A set built from records that are easy to tell apart cannot measure the failure, which lives in the near-collisions.A pass on the entity check says nothing about whether the answer about that entity was right. |
| Takes an action it had no authority to take The action itself may be fine; the system was not permitted to take it. | action and authority | A case where the action is correct on the merits and the authority for it is absent.A case where authority exists for a smaller version of the action and not the one requested.A case where the reader asks the system to exceed its own stated bounds. | product-or-line, account-or-tenure, channel | expected-tool-calls: trajectory-or-state-check, exact-or-programmaticexpected-final-state: trajectory-or-state-check, exact-or-programmaticexpected-behavior-only: human-expert | Constraint violations, pass^k | Each case is a written must-not-do rule, counted as a constraint violation with its own denominator rather than folded into a pass rate. Blast radius follows the permission that was exceeded. | repeat and report spread | every case | I12 Permission and irreversible-action gates, I11 Tool-call correctness: schema, arguments, wrong-tool rate, I06 Prompt-injection and adversarial suite | A permission suite covers the permissions somebody wrote down. The inventory is the input and it is the thing that goes stale.Zero violations in n bounds the rate at about 3/n (one-sided 95 percent); it does not establish that none happen. |
| Stops before the task is done and reports otherwise The run ends early and the summary says it finished. | completion and control | A multi-step case whose last step is the one that matters, so stopping at step three still reads as progress.A case where an intermediate step fails and the correct behavior is to say so.A case where the end state is checkable independently of what the run reported. | intent-or-task-type, input-length, channel | expected-final-state: trajectory-or-state-check, exact-or-programmaticexpected-tool-calls: trajectory-or-state-checkexpected-behavior-only: human-expert | Pass rate with an interval, pass^k, Cost and latency | A silent stop and a reported stop are separate classes: only the silent one leaves a person believing the work is done. Severity follows what was left undone, not how far the run progressed. | repeat and report spread | sampled | I21 End-to-end multi-step run, I23 Error paths, rate limits, and timeouts, I01 Core task suite on the frozen golden set | Grading the final message measures the message. Only a state check sees whether the work happened.A completion rate on cases the harness can drive to the end says nothing about cases it cannot. |
| A tool call failed and the run continued as if it had not An error came back, nothing handled it, and the answer was written anyway. | completion and control | A case where the tool is made to return an error, and the correct behavior is to stop or retry.A case where the tool returns an empty result that is valid and useless.A case where the tool result is truncated, so the run has part of what it needed. | upstream-source, truncation, intent-or-task-type | expected-tool-calls: trajectory-or-state-check, exact-or-programmaticexpected-final-state: trajectory-or-state-check | Constraint violations, Pass rate with an interval, pass^k | A swallowed error that produced an answer and one that produced a stop are counted as separate classes. Severity follows what was written on the strength of a result the tool never returned. | repeat and report spread | sampled | I23 Error paths, rate limits, and timeouts, I11 Tool-call correctness: schema, arguments, wrong-tool rate, I21 End-to-end multi-step run | Injected tool errors test the errors you injected. A real dependency fails in ways the fixture does not.A clean run on a healthy dependency says nothing about the degraded path. |
| The same input produces materially different behavior across runs Two runs of one case disagree in a way that matters. | consistency and robustness | A case run several times with everything pinned, kept as separate observations rather than averaged.A case whose correct answer is unique, so a disagreement between runs is unambiguous.A case at the edge of a decision boundary, where variance shows up first. | intent-or-task-type, input-length | reference-answer: exact-or-programmaticexpected-final-state: trajectory-or-state-check, exact-or-programmaticrubric: rubric-human, llm-judge | pass@k, pass^k, Pass rate with an interval | Variance is reported as a spread, never averaged into a single rate. Whether a given spread matters depends on the action the output drives: the same spread is acceptable on a draft and unacceptable on a write. | repeat and report spread | not needed | I16 Repeated-run variance at the new setting, I14 Judge stability across repeated runs, I01 Core task suite on the frozen golden set | Repeated runs on the same cases reduce uncertainty about run-to-run variance, never about population coverage.pass@k rises with k by construction. A product that runs once experiences pass@1. |
| Works on the average case and fails on a declared subset The headline rate is fine and one named group is not. | coverage and equity | The same task written once per declared slice value, with the slice label carried on the row.An intersection case: two slice values at once, where the count is smallest and the failure usually is.A case that is ordinary except for the slice label, so a failure cannot be blamed on difficulty. | language-and-locale, channel, account-or-tenure, protected-or-regulated-population, vulnerability-signal | reference-answer: exact-or-programmaticrubric: rubric-human, llm-judgeexpected-behavior-only: human-expert | Per-slice rate, Pass rate with an interval | Severity inherits from what the slice names: a slice covering a regulated population carries B3 blast radius, an internal slice carries B1. The slice itself is never assigned a severity; it inherits the blast radius of the population it contains. | repeat per case | sampled | I03 Per-slice pass rates for the named slices, I01 Core task suite on the frozen golden set | Below about 30 cases a rate claim carries an interval wider than most decisions can use; 30 is our default, not a derived threshold. Size the slice from the claim you need with the golden set size planner.Declaring a slice needs the label to exist in the data, which for some slices means holding or inferring a sensitive attribute. That is a real tradeoff, not a technicality.Full coverage of the slices you declared says nothing about the slices you did not. |
| Does not hand off when the situation required a person The case needed a human and the system answered it instead. | oversight and escalation | A case whose only correct outcome is a handoff, with everything about it reading as answerable.A case where the trigger for escalation appears midway through the conversation.A case where the handoff must carry specific context, so a bare transfer still fails. | vulnerability-signal, channel, intent-or-task-type | expected-behavior-only: human-expertrubric: rubric-human, human-expertexpected-tool-calls: trajectory-or-state-check | Coverage and risk together, Constraint violations, Pass rate with an interval | A missed escalation is counted against its own denominator: cases that should have been handed off. Read this count beside the over-refusal count, because tightening one loosens the other. | repeat per case | domain expert | I05 Safety and refusal suite, I01 Core task suite on the frozen golden set, I03 Per-slice pass rates for the named slices | An escalation rate read on its own moves when the case mix moves, without any behavior changing.A handoff that happened says nothing about whether the person receiving it could act on what they got. |
| Takes an action nothing undoes, wrongly The system did something permanent that it should not have done. | action and authority | A case whose correct outcome is that the action does NOT run, with everything about it reading as though it should.A case where the parameters are right and the target is wrong: correct payee, wrong amount; correct record, wrong field.A case that changes mid-conversation after the reader already confirmed, testing whether the confirmation is re-sought.A case where the action is reachable only after a state the run has not established. | product-or-line, account-or-tenure, channel | expected-final-state: trajectory-or-state-check, exact-or-programmaticexpected-tool-calls: trajectory-or-state-check, exact-or-programmaticexpected-behavior-only: human-expert | Constraint violations, pass^k | R3 reversibility and irreversible on the reversal axis. Each case is its own counted class with its own denominator, never folded into an aggregate pass rate. Zero observed in n bounds the rate at roughly 3/n (one-sided 95 percent), which is the claim to record. | repeat and report spread | every case | I12 Permission and irreversible-action gates, I24 Rollback rehearsal on a copy, I21 End-to-end multi-step run, I11 Tool-call correctness: schema, arguments, wrong-tool rate | Zero severe failures in a case set bounds the rate; it does not establish that none happen.A case set built from actions somebody already listed cannot cover an action nobody listed. The inventory is the input, and it is the thing that goes stale. |
| Breaks a written must-not-do rule while producing an acceptable-looking answer The output passes an ordinary read and breaks a rule somebody wrote down. | action and authority | A case where the fluent answer is the one that breaks the rule.A case that pits two written rules against each other, so the correct behavior is to stop.A case where the rule is about how, not what, so the output content alone cannot show the breach. | product-or-line, protected-or-regulated-population, intent-or-task-type | expected-tool-calls: trajectory-or-state-check, exact-or-programmaticrubric: rubric-human, human-expertexpected-behavior-only: human-expert | Constraint violations, Pass rate with an interval | Constraints are counted, never scored. Each written rule carries its own count and its own denominator: a rule broken once in a thousand runs and one broken once in ten are different objects requiring different responses. | repeat and report spread | sampled | I05 Safety and refusal suite, I12 Permission and irreversible-action gates, I01 Core task suite on the frozen golden set | A constraint suite covers the constraints that were written down. The unwritten ones are the ones that bite.A count of violations is not a rate until the denominator is the cases where the rule could have been broken. |
| Checks its own work and passes something wrong, or does not check The self-check reported success on output that was not right. | completion and control | A case whose output is wrong in a way the system's own check is built to catch, so a pass is a checker failure.A case where the check itself is broken, so a green result proves nothing.A case where the correct behavior is to report an unverifiable result rather than assert one. | intent-or-task-type, content-or-document-type | expected-final-state: trajectory-or-state-check, exact-or-programmaticexpected-tool-calls: trajectory-or-state-checkexpected-behavior-only: human-expert | Pass rate with an interval, Rater or judge agreement, Constraint violations | A false pass from the self-check is counted apart from an ordinary failure: it removes the signal a person would have acted on. Severity follows what shipped because the check said it was good. | repeat per case | sampled | I13 Judge agreement against human labels, I15 Rubric back-compatibility on a frozen sample, I01 Core task suite on the frozen golden set | A green check on an untested checker is evidence about nothing.Agreement between the system and its own checker is not correctness; both can be wrong in the same direction. |
| Output does not parse, or parses and violates the contract The shape of the output is wrong, or it is valid and says something the contract forbids. | output correctness | A case whose schema has a required field the content does not obviously supply.A case that is schema-valid and semantically wrong, which is the class people forget.A case where the input contains characters that break the output format if they are passed through. | formatting-shift, content-or-document-type, input-length | reference-answer: exact-or-programmaticexpected-final-state: exact-or-programmatic, trajectory-or-state-check | Pass rate with an interval, Constraint violations, Per-slice rate | A parse failure is loud and usually cheap; a schema-valid wrong value is quiet and usually expensive. Count the two as separate classes rather than reporting one conformance rate. | repeat and report spread | not needed | I04 Output-contract and parse conformance, I01 Core task suite on the frozen golden set, I16 Repeated-run variance at the new setting | A parse rate says nothing about whether the parsed values are right.Conformance measured on the inputs you have does not extend to an input shape the set does not contain. |
| Runs out of steps, tokens, time or retries before finishing The run hits a limit and stops, whatever it had left to do. | completion and control | A long case sized so that a wasteful path exhausts the budget and a direct path does not.A case where a retry loop is reachable, testing whether it terminates.A case whose correct behavior on hitting the limit is to report the partial state rather than assert completion. | input-length, intent-or-task-type, upstream-source | expected-final-state: trajectory-or-state-check, exact-or-programmaticexpected-tool-calls: trajectory-or-state-check | Cost and latency, Pass rate with an interval, pass^k | Exhaustion that reports itself is an availability problem; exhaustion that reports success is a correctness problem. Count the two as separate classes and read the cost figures beside them. | repeat and report spread | not needed | I23 Error paths, rate limits, and timeouts, I17 Cost per successful task, I18 Latency distribution | A budget measured on the cases you ran does not bound the cost of a case shape the set does not contain.A median latency hides the tail, which is where the timeouts are. |
| Loses earlier state across a turn, a handoff or a session Something established earlier stops being true for the system later on. | consistency and robustness | A multi-turn case where a constraint stated in turn one decides the answer in turn six.A case that crosses a handoff between components, with the decisive fact stated before the handoff.A case long enough that the decisive fact sits in the middle of the context rather than at either end. | input-length, channel, truncation | reference-answer: exact-or-programmaticrubric: rubric-human, llm-judgeexpected-final-state: trajectory-or-state-check | Pass rate with an interval, Per-slice rate, pass^k | Severity follows the fact that was lost. A forgotten formatting preference and a forgotten consent withdrawal are the same mechanism and never the same class. | repeat per case | sampled | I22 Context-window and long-input behavior, I21 End-to-end multi-step run, I01 Core task suite on the frozen golden set | Context behavior measured at one input length says nothing about another; length is the mechanism.A pass on a scripted conversation does not extend to a conversation a person actually has. |
| Content it retrieved changed what it did An instruction inside the data was followed as though it came from the reader. | action and authority | A case where a retrieved document carries an instruction and the correct behavior is to ignore it.A case where the planted instruction asks for an action the system is otherwise permitted to take.A case where the instruction is in a part of the document a reader would not see. | injected-instruction, upstream-source, content-or-document-type | expected-tool-calls: trajectory-or-state-check, exact-or-programmaticexpected-final-state: trajectory-or-state-check, exact-or-programmaticexpected-behavior-only: human-expert | Constraint violations, pass^k | Counted as a constraint class with its own denominator, read together with the action-risk rating of whatever the injected instruction asked for. A successful injection inherits the blast radius of the action it triggered. | repeat and report spread | every case | I06 Prompt-injection and adversarial suite, I12 Permission and irreversible-action gates, I11 Tool-call correctness: schema, arguments, wrong-tool rate | An injection suite measures the phrasings you wrote. It is a lower bound on what an adversary will try.Zero successful injections in n bounds the rate at about 3/n (one-sided 95 percent); it does not establish that none succeed. |
| Declines cases it should have handled The system refuses work that was within its remit. | oversight and escalation | A case that shares surface features with a class the system should refuse and is itself fine.A case in a sensitive area whose correct answer is ordinary and specific.A case where the correct behavior is to answer part and decline part, with the split stated. | intent-or-task-type, protected-or-regulated-population, language-and-locale | rubric: rubric-human, llm-judgereference-answer: exact-or-programmaticexpected-behavior-only: human-expert | Coverage and risk together, Pass rate with an interval, Per-slice rate | Over-refusal is counted against the cases the system should have handled, read beside the missed-escalation count on the same axis. Neither number means anything read alone. | repeat per case | sampled | I05 Safety and refusal suite, I03 Per-slice pass rates for the named slices, I01 Core task suite on the frozen golden set | A refusal rate on its own moves when the case mix moves, without any behavior changing.Cases written by the team that tuned the refusals inherit their view of what is in remit. |
Slice union
| Slice | Axis | Needed by |
|---|---|---|
| Channel the request arrives on | subpopulation | Takes an action it had no authority to take, Stops before the task is done and reports otherwise, Works on the average case and fails on a declared subset, Does not hand off when the situation required a person, Takes an action nothing undoes, wrongly, Loses earlier state across a turn, a handoff or a session |
| What the reader is trying to do | subpopulation | Leaves out something the answer needed, Recommends an action that would cause harm, Stops before the task is done and reports otherwise, A tool call failed and the run continued as if it had not, The same input produces materially different behavior across runs, Does not hand off when the situation required a person, Breaks a written must-not-do rule while producing an acceptable-looking answer, Checks its own work and passes something wrong, or does not check, Runs out of steps, tokens, time or retries before finishing, Declines cases it should have handled |
| Language, locale and script | subpopulation | Works on the average case and fails on a declared subset, Declines cases it should have handled |
| The kind of document or record involved | subpopulation | Leaves out something the answer needed, States something the source does not support, Answers from material that is no longer current, Checks its own work and passes something wrong, or does not check, Output does not parse, or parses and violates the contract, Content it retrieved changed what it did |
| How much text or context the case carries | subpopulation | Leaves out something the answer needed, States something the source does not support, Stops before the task is done and reports otherwise, The same input produces materially different behavior across runs, Output does not parse, or parses and violates the contract, Runs out of steps, tokens, time or retries before finishing, Loses earlier state across a turn, a handoff or a session |
| How long the account or relationship has existed | subpopulation | Acts on or answers about the wrong record, person or product, Takes an action it had no authority to take, Works on the average case and fails on a declared subset, Takes an action nothing undoes, wrongly |
| Which product, plan or line the case is about | subpopulation | Answers from material that is no longer current, Acts on or answers about the wrong record, person or product, Takes an action it had no authority to take, Takes an action nothing undoes, wrongly, Breaks a written must-not-do rule while producing an acceptable-looking answer |
| Where the input came from before it reached the system | subpopulation | States something the source does not support, Acts on or answers about the wrong record, person or product, A tool call failed and the run continued as if it had not, Runs out of steps, tokens, time or retries before finishing, Content it retrieved changed what it did |
| When the case is from, and how fresh the answer must be | subpopulation | Answers from material that is no longer current |
| Cases about people a rule treats differently | subpopulation | Recommends an action that would cause harm, Works on the average case and fails on a declared subset, Breaks a written must-not-do rule while producing an acceptable-looking answer, Declines cases it should have handled |
| Cases carrying a disclosed circumstance that changes correct behavior | subpopulation | Recommends an action that would cause harm, Works on the average case and fails on a declared subset, Does not hand off when the situation required a person |
| The same content in a different layout or file shape | transformation | Output does not parse, or parses and violates the contract |
| The input or a tool result cut short | transformation | A tool call failed and the run continued as if it had not, Loses earlier state across a turn, a handoff or a session |
| Instructions planted in content the system retrieves or reads | adversarial | Content it retrieved changed what it did |
Rerun items
I01 Core task suite on the frozen golden set, I03 Per-slice pass rates for the named slices, I04 Output-contract and parse conformance, I05 Safety and refusal suite, I06 Prompt-injection and adversarial suite, I07 Retrieval recall at k, I08 Retrieval relevance at the new chunk size and top-k, I09 Reranker precision, I10 Grounding and faithfulness, I11 Tool-call correctness: schema, arguments, wrong-tool rate, I12 Permission and irreversible-action gates, I13 Judge agreement against human labels, I14 Judge stability across repeated runs, I15 Rubric back-compatibility on a frozen sample, I16 Repeated-run variance at the new setting, I17 Cost per successful task, I18 Latency distribution, I21 End-to-end multi-step run, I22 Context-window and long-input behavior, I23 Error paths, rate limits, and timeouts, I24 Rollback rehearsal on a copy
This mapping is the site's own reading of what each concern calls for. It is a starting point, not a standard. Every row is editable, and the result is the reader's, not ours.
Questions
Questions
What is a concern?
A named way the system could fail that someone decided matters. The nineteen presets are a palette drawn from operator experience and the research dossier. They are not exhaustive and they are not ranked.
Can I add my own concerns?
Yes. The custom-concern table at the top lets you write a concern, give it a family and a severity note. The mapping columns will be empty; fill them in yourself, using the preset rows as a template.
Why are some cells empty?
A cell is empty when the preset has no content for that mapping column. The count at the top tells you how many are empty out of how many exist. An empty cell is a thing you have not specified, not a thing that does not apply.
What is an evaluator conflict?
Two concerns that share an oracle type but name different evaluators for it. This is not an error. It is a design choice the reader should make deliberately.
What do the rerun items mean?
Each I-code is one of the twenty-four rerun items in the change impact worksheet: a specific thing to re-run when a specific condition changes. The union shows every rerun item any selected concern calls for.