LatentEval

For builders

Golden set blueprint: coding agents

Seven case families for coding agents: must-flip and must-not-break tests on separate denominators, solution leakage measured, and the test suite treated as the evaluator it is.

For builders

In brief

3 POINTS
  • Your test suite is the evaluator, so an unaudited suite is an unvalidated evaluator producing numbers about a grader nobody checked.
  • You run must-not-break tests alongside must-flip tests, because a patch that turns one test green while breaking three others is a net loss.
  • When the issue text contains the fix, copying it and watching tests pass measures clipboard operation, not debugging capability.

Why this family’s correctness is different

The test suite is the evaluator, not the output text. Most coding-agent evals treat passing tests as the finish line and stop there: the agent either fixed the bug or it did not. But the test suite doing the grading is itself an instrument, and an unaudited instrument validates nothing. A test that asserts the current output rather than the specified behavior passes everything and proves nothing. A patch that makes the target test green while breaking three others is a net loss the harness calls a win. And when the issue text already contains the fix, copying it and watching the tests go green measures clipboard operation, not debugging capability. These are not edge cases. They are the three most common ways a coding-agent eval inflates its own numbers, and the seven case families in this blueprint exist to separate them.

The decisive mistake in a coding-agent run usually lands in the first file the agent edits and stays invisible for many subsequent steps. That is why trajectory matters alongside final state: an end-state check that only asks “do the tests pass?” misses the patch that changed the wrong abstraction layer, fixed the symptom, and left the root cause intact. Solution leakage inside the issue text is a different problem; it inflates success without inflating capability, and the gap between a leaked and an unleaked variant of the same bug is the inflation estimate your eval should report.

This blueprint rests on research published on this site. The trajectory study of how coding agents fail mapped the failure shapes that recur across runs (wrong first file, cascading edits, scope creep) and showed that final-state pass rates alone hide most of them. The agent evaluation framework laid out the structure an eval needs to survive its own optimism: oracle types, evaluator pairings, severity classes, and the coverage skeleton that tells you what you have not tested. These are the evidence this blueprint is built on.

Case families

A case family is a class of coding task that tests one property of the agent. Each family names what it tests, the oracle that decides correctness, and the evaluator that applies it.

These counts seed a starter set; none of them supports a rate claim on its own. The coverage matrix floors set the claimable and readable thresholds.

A set built from this pack carries several oracle types across its families; run the evaluator navigator once per family and write acceptance criteria per family.

TABLEShow full table (7 rows)Showing full table (7 rows)
Case familyWhat it testsExpected behaviorUnacceptable behaviorOracle typeEvaluatorReviewer expertiseSeverity treatmentStarter rows to writeWhat this count licenses
must-flip-testWhether the agent produces a change that makes the specified failing test pass, verified by running the test suite before and after.The named test fails before the patch and passes after it, with no manual intervention between.The test still fails after the patch; the test was deleted or weakened instead of the code being fixed; the test passes only because its assertion was changed to match the wrong output.expected-final-stateexact-or-programmatic(none)R1 where the patch is on a branch and revertible, R2 where it has been merged to a shared branch. Counted as its own class so the must-flip rate is never folded into a must-not-break rate.15 - 30At 30, worst-case 95% interval: +/-16.8 points
must-not-breakWhether the full test suite that passed before the patch still passes after it, catching regressions the agent introduced.Every test that passed before the patch still passes after it. No new test failures appear.A previously passing test now fails; a previously passing test was deleted or skipped to hide the regression.expected-final-stateexact-or-programmatic(none)R1 where the patch is on a branch and revertible, R2 where it has been merged to a shared branch. B1 for a broken unit test, B2 for a broken integration test that gates deployment. Counted as its own class, never averaged with must-flip.15 - 30At 30, worst-case 95% interval: +/-16.8 points
solution-leakageWhether the agent actually diagnosed the problem or merely extracted the answer that was already present in the issue description, inflating success without demonstrating capability.When the issue text is stripped of its embedded solution, the agent still produces a correct fix through its own reasoning.The agent succeeds only when the fix is spelled out in the issue and fails when it is removed; the agent copies the suggested code verbatim without verifying it against the test suite.expected-final-stateexact-or-programmatic(none)Not a production severity issue but a measurement validity issue. A leaked solution inflates pass rate without inflating capability, so cases with leakage are counted in their own denominator and the gap between leaked and unleaked is the inflation estimate.8 - 15At 15, worst-case 95% interval: +/-22.6 points
early-decisive-errorWhether the agent recovers or compounds a wrong first move. The decisive mistake lands early, looks plausible, and stays invisible for many subsequent steps.The agent either gets the first change right, or recognizes the error when later steps contradict it and backtracks.The first change is wrong and every subsequent file is built on the wrong assumption; the agent produces a chain of internally consistent but globally incorrect changes without re-examining the premise.expected-final-statetrajectory-or-state-check(none)R1 where the patch is on a branch and revertible, R2 where it has been merged. B1 for a single-file cascade, B2 for a cascade that rewrites the module boundary. The trajectory check must record the point of divergence, not only the final state.6 - 12At 12, worst-case 95% interval: +/-24.6 points
multi-file-coordinationWhether the agent keeps type signatures, import paths, configuration, and behavioral contracts consistent when the change touches more than one file.All changed files compile, pass type-checking, and pass tests together. No file references a symbol, path, or contract that another changed file no longer provides.A renamed function is called by its old name in another file; a type signature was updated in the definition but not in the callers; a configuration value was changed in one place and left stale in another.expected-final-stateexact-or-programmatic(none)R1 where the patch is on a branch and revertible, R2 where it has been merged. B1 where the inconsistency is caught by the compiler, B2 where it passes compilation and fails at runtime. The count separates compile-time from runtime inconsistencies because they carry different blast radii.10 - 20At 20, worst-case 95% interval: +/-20.1 points
test-suite-validityWhether tests the agent authored actually test what they claim to test. An unaudited test suite is an unvalidated evaluator: a test that asserts the wrong thing passes quietly and validates nothing.Each test the agent wrote has a clear assertion tied to the requirement, tests the behavior rather than the implementation detail, and fails when the behavior is broken.A test that always passes regardless of the code under test; a test whose assertion matches the current output rather than the specified behavior; a test that mocks the component it is supposed to be testing.rubricrubric-humanAn engineer who writes tests in the language under review and can distinguish a behavior assertion from a snapshot of the current output.A bad test that passes is worse than a missing test, because it actively hides the gap. Counted as its own class. R1 where the test can be rewritten on a branch. B1 for a unit test, B2 for a test that gates deployment.8 - 15At 15, worst-case 95% interval: +/-22.6 points
scope-containmentWhether the agent limits its changes to what was asked for, or whether it refactors, reformats, or “improves” code that was not part of the task.The diff contains only changes that are necessary to fulfill the stated task. Unchanged files remain unchanged. Formatting-only changes do not appear unless formatting was the task.Files outside the task scope were modified; unrelated refactoring was performed alongside the requested fix; whitespace or style changes were made to lines not involved in the fix.expected-final-statetrajectory-or-state-check(none)R0 where the out-of-scope change is cosmetic and revertible. R1 where it changes behavior in code the task did not mention. R2 where it changes a shared interface or public API. Counted as its own class because an agent that fixes the bug and also breaks the neighbor passes the first check and fails this one.8 - 15At 15, worst-case 95% interval: +/-22.6 points

The first two families, must-flip-test and must-not-break, belong together but are counted separately. A patch that makes the target test pass and breaks two others is a net negative, but an eval that folds both into one “tests pass” aggregate reports it as a success. Keeping them in separate denominators is what makes the regression rate visible.

Slice families

A slice family is a dimension you cut your results across after the run. Every case belongs to one or more slices, and the slice tells you whether performance is uniform or depends on something invisible in the aggregate.

TABLEShow full table (5 rows)Showing full table (5 rows)
SliceWhy it mattersSensitive
input-lengthA single-file fix and a cross-module refactor exercise different failure paths. Agents that succeed on small diffs may fail when the relevant context spans many files.No
programming-languageCapability varies by language. An agent that fixes Python reliably may hallucinate Go idioms, and a pass rate across languages hides which ones are holding.No
content-or-document-typeLibrary code, application code, test code, build configuration, and infrastructure-as-code each have different correctness criteria and different blast radii on failure.No
upstream-sourceA typed issue with a stack trace, a generated issue from a CI failure, and a vague feature request each carry different amounts of solution leakage and different ambiguity.No
intent-or-task-typeBug fixes, feature additions, refactors, test authoring, and dependency upgrades each exercise different agent capabilities and carry different failure shapes.No

None of these slices are marked sensitive because they describe the code and the task, not the person or the context around it.

Coverage matrix skeleton

The coverage skeleton lists the case-family-plus-slice intersections where an empty cell means you are flying blind on something that matters. It is not a test plan. It is the minimum set of cells whose absence leaves a gap you should know about.

The twelve entries for this pack:

  1. must-flip-test x programming-language. You would not know which programming languages the agent can actually fix bugs in and which it cannot.
  2. must-flip-test x intent-or-task-type. You would not know whether the agent handles bug fixes differently from refactors or feature work.
  3. must-not-break x input-length. You would not know whether regression rate increases with the size of the change.
  4. must-not-break (any slice). You would not know how often the agent introduces regressions while making the target fix.
  5. solution-leakage x upstream-source. You would not know how much of the measured pass rate comes from the solution already being in the issue text.
  6. early-decisive-error (any slice). You would not know whether the agent recovers from a wrong first move or compounds it across the rest of the change.
  7. multi-file-coordination x input-length. You would not know whether cross-file consistency holds as the number of touched files increases.
  8. multi-file-coordination x programming-language. You would not know whether the agent keeps types and imports consistent across files in each language it supports.
  9. test-suite-validity x content-or-document-type. You would not know whether agent-authored tests for different code types actually test what they claim to.
  10. test-suite-validity (any slice). You would not know whether the test suite the agent produced is itself a valid evaluator.
  11. scope-containment x intent-or-task-type. You would not know which task types trigger out-of-scope changes and which stay contained.
  12. scope-containment (any slice). You would not know whether the agent limits itself to the requested change or edits unrelated code.

Entries 4, 6, 10, and 12 have no slice because the property they test is family-level: regression rate, cascade recovery, test validity, and scope discipline do not depend on a dimension. They depend on whether you tested them at all.

Evaluator and annotation guidance

Each oracle-evaluator pairing carries its own validity evidence: the checks you need to run on the grading instrument itself before you trust the numbers it produces.

expected-final-state / exact-or-programmatic. This covers must-flip-test, must-not-break, solution-leakage, and multi-file-coordination: every family where the test suite is the oracle. The test suite is run by the harness, not by the agent, so the agent cannot alter what is measured. The fixture is reset to the pre-patch state between runs, so one run cannot grade another run’s leftovers. Both must-flip and must-not-break verdicts are collected from the same test run, so a fix that breaks something is caught in the same pass. The metric families that apply are pass-rate, pass^k, and constraint-violation-count.

expected-final-state / trajectory-or-state-check. This covers early-decisive-error and scope-containment: families where the final state alone does not tell the whole story. The trajectory is recorded by the harness from the agent’s tool calls, not reconstructed from the agent’s narrative about what it did. The diff is computed against the known pre-patch state, so an out-of-scope change cannot hide inside a large patch. A known-bad trajectory (a wrong first file, an out-of-scope edit) is fed to the checker and it flags it. The metric families are constraint-violation-count and pass-rate.

rubric / rubric-human. This covers test-suite-validity, the one family where a programmatic check is not enough. The reviewer can write tests in the language under review and has stated criteria for what makes a test valid. Each rubric criterion is exercised against at least one known-good and one known-bad test, confirming the criterion discriminates. At least two reviewers grade a sample and inter-rater agreement is reported. The metric families are agreement, pass-rate, and constraint-violation-count.

The common thread: every pairing requires you to validate the instrument, not just the output. A test suite needs a clean fixture. A trajectory checker needs a known-bad case. A rubric needs inter-rater agreement. Without that validation, the pass rate is a number about a grader nobody tested.

Starter rows

These rows are starter examples showing what a case in this family looks like. They are not validated gold-standard answers. Before any of them grades a real output, someone who knows your system needs to confirm or rewrite the expected behavior in each one.

The pack ships eleven starter rows, ca-001 through ca-011, that illustrate the seven case families:

ca-001 and ca-002 are a solution-leakage pair. ca-001 gives the agent a failing auth test and tells it the root cause (the operator uses <= instead of <). ca-002 is the same bug with the diagnosis removed. Together they measure the leakage delta: how much of the pass rate on ca-001 comes from the solution already being in the issue text rather than from the agent’s own reasoning. Both are must-flip-test cases, but when run as a pair they answer the solution-leakage question.

ca-003 is a multi-file-coordination case. The agent extracts email validation from a service class into a standalone utility and updates all callers. The test is whether type signatures, import paths, and behavior stay consistent across the service, the new utility file, and three calling modules.

ca-004 is a feature addition. The agent adds a --dry-run flag to a deploy CLI command and writes tests for both modes. This row checks whether the feature works: the flag is accepted, no side effects occur in dry-run mode, and the tests pass.

ca-011 is the test-suite-validity companion to ca-004. Same task, same output, but the oracle is now a rubric applied by a human reviewer who audits the tests the agent authored. A test that asserts the current output rather than the specified behavior validates nothing, and this row catches it.

ca-005 is an early-decisive-error case. A NullPointerException in a refund processor can be fixed in the right place (the processor) or the wrong place (the data model). The wrong first move (changing the model so it never returns null) looks plausible, passes the immediate test, and breaks the contract for every other caller. The trajectory check records whether the agent touched the model at all.

ca-006 is a multi-file-coordination case disguised as a dependency upgrade. Upgrading a logging library from v2 to v3 requires renaming every call site and converting the config file format. The test is whether every call site was caught and whether the config file validates against the new schema.

ca-007 is a concurrency bug fix with a test-authoring requirement. The agent fixes a race condition in a connection pool and writes a regression test. The regression test is itself tested: it must fail on the pre-fix code and pass on the post-fix code. A regression test that passes on both proves it cannot catch the bug.

ca-008 is a scope-containment case. A CSS layout bug requires a stylesheet-only fix. The agent may be tempted to restructure the HTML or add JavaScript. The trajectory check confirms the diff touches only CSS files.

ca-009 is a test-suite-validity case. The agent writes unit tests for a discount function, covering edge cases. The rubric checks whether each test asserts the business rule or merely snapshots the current return value, and whether each test fails when the behavior it claims to guard is broken.

ca-010 is a scope-containment case with an explicit prohibition. The agent ports a Python data pipeline to a new API, with the constraint that no test file may be modified. The test checks both the migration (no old API calls remain) and the scope (git diff shows no changes in the test directory).

Failure-to-case mapping

When something goes wrong in production, the case families tell you where to look.

The agent edits the test assertion instead of fixing the bug. That is must-flip-test. The test flipped, but only because the agent changed what it expected. Your golden set needs cases where the expected behavior is defined by the original assertion, and the grader checks that the assertion was not weakened.

The agent fixes the target bug but breaks unrelated tests. That is must-not-break. The must-flip rate looked fine because you were only watching the target test. The regression rate is the separate measurement that catches this, which is why it lives in its own denominator.

The issue text contains the fix and the agent just copies it. That is solution-leakage. Run the same case with and without the diagnosis in the issue text. The gap between the two pass rates is the inflation estimate, and it tells you how much of your measured capability is actually a clipboard operation.

The agent edits the wrong file first and everything after it is built on a wrong assumption. That is early-decisive-error. The final state may even pass if the wrong assumption happens to produce code that satisfies the tests, which is why the trajectory check records the point of divergence, not just the outcome.

A renamed function is still called by its old name in another file. That is multi-file-coordination. The project compiles only by accident (the old name still exists as a different function), or it does not compile at all. The case tests whether the agent tracked the rename across every caller.

The agent wrote tests that always pass regardless of the code under test. That is test-suite-validity. A test that cannot fail is worse than a missing test, because it actively hides the gap it was supposed to guard. The rubric asks whether each test fails when the behavior it claims to test is broken.

The agent reformatted files that had nothing to do with the task. That is scope-containment. The fix works, but the diff is three lines of fix and two hundred lines of whitespace changes. The trajectory check separates the task-related changes from the out-of-scope ones.

Downloads

The eleven starter rows are downloadable as CSV and JSONL:

Both formats carry the same data. The CSV is easier to open in a spreadsheet; the JSONL preserves the nested case profile without flattening. This family’s files are published as coding-agent.csv and coding-agent.jsonl. To edit them in the browser and add your own, use the blueprint builder.

Which instruments apply

Each evaluator guidance entry maps to instruments in the LatentEval toolkit.

exact-or-programmatic evaluators (must-flip-test, must-not-break, solution-leakage, multi-file-coordination): use the pass-rate confidence interval calculator to put an interval on the rate, and the reliability-at-k estimator to measure how often the agent gets it right on the first try versus needing multiple attempts. The metric families to track are pass-rate, pass^k, and constraint-violation-count. Pass^k matters more than pass-rate here because a coding agent that succeeds on attempt three after failing twice is a different instrument than one that succeeds on attempt one, and a single pass rate hides the difference.

trajectory-or-state-check evaluators (early-decisive-error, scope-containment): use the pass-rate confidence interval calculator for the overall rate and the agent action risk matrix to classify the trajectory failures by severity and blast radius. The metric families are constraint-violation-count and pass-rate. The constraint-violation-count is the more informative one because the question is not whether the agent usually stays in scope but what happens when it does not.

rubric-human evaluators (test-suite-validity): use the inter-rater reliability calculator to measure whether your rubric produces consistent verdicts across reviewers. If agreement is low, the rubric needs rewriting, not more raters. The metric families are agreement, pass-rate, and constraint-violation-count.

Across all families: use the eval coverage matrix for the grid of which cells are populated and which are empty, and the risk-to-test mapper for mapping from a named concern to the case families that test for it.

Companion pack guides

  • Research agents covers citation verification and synthesis fidelity, where the failure surface is fabricated evidence rather than broken code.
  • RAG and knowledge assistants covers retrieval-grounded generation, where the shared concern is whether the system grounds its output in the sources it was given.