INSTRUMENT | Eval statistics
Evaluation Coverage Matrix: Which Cells Have No Cases
Put what you test on one axis and what you test it across on the other. The grid shows which cells have cases, which are empty, which are too small to carry an interval, and what each gap costs you.
Declare what you are testing on one axis and what you are testing it across on another, then see which cells are empty, which are too small to read, and what an empty cell means you do not know. To measure an existing dataset's actual distribution, use the slice and class balance analyzer. Grounded in the testing checklist and the rerun evidence. The risk-weight and persona axes are not built yet. Risk weights come from the action risk matrix and the cost a miss carries, computed in the acceptable error rate calculator.
2 of 9 cells are empty.
| Wrong answer | Refusal or timeout | Safety violation | |
|---|---|---|---|
| Happy path | 120 ±8.8% (max) Claimable | 40 ±14.8% (max) Claimable | 0 Empty Whether the system works under normal conditions at all |
| Edge input | 25 Below claim floor | 8 Below read floor | 3 Below read floor |
| Adversarial | 0 Empty Whether the system resists deliberately hostile input | 15 Below read floor | 30 ±16.8% (max) Claimable |
A count of empty cells and their positions. No coverage percentage, no grid score: two integers, read as two integers.
cell status = the count against the read floor, the claim floor and the interval it would carryHow?
How this is calculated
Each cell in the grid gets exactly one status. A cell with zero cases is empty. A cell with 1 to 15 cases is below the read floor: too few to carry even a wide confidence interval (16 is the smallest n whose Wald 95% half-width at p = 0.5 is within ±25 points (n ≥ 4z²); cells print the Wilson half-width, which is narrower). A cell with 16 to 29 cases carries a readable interval but a wide one. Whether it supports your claim depends on the rate you want to bound: see slice planning, where n is derived from the target rate rather than a fixed number. Below about 30 cases a rate claim carries an interval wider than most decisions can use; 30 is our default, not a derived threshold. Size the slice from the claim you need with the golden set size planner. A cell with 30 or more cases is claimable, and the widest Wilson half-width the cell could carry (max, computed at p = 0.5) is printed beside the count; a rate away from 0.5 gives a narrower interval.
In slice-plan mode, the rule-of-three sizing shows how many cases you need to claim a rate below a target with zero observed failures. This is the one-sided 95% Clopper-Pearson upper limit at zero failures: n ≥ 3/r (the mnemonic) beside the exact bound 1 − 0.051/n. The acceptance criteria builder sizes the same claim with the two-sided bound, which returns a larger n for the same r (r = 0.001: 3000 here, 3688 there; r = 0.01: 300 here, 368 there). Intersection sparsity multiplies the shares of two slice families to show the expected count at the crossing, which is where the count is smallest and the failure usually is. The product assumes the two families are independent; a correlated pair will be sparser or denser than the product, so treat the figure as a planning floor, not a forecast.
Formula: cell status = the count against the read floor, the claim floor and the interval it would carry
Questions
Questions
What do the four modes cover?
Scenario mode crosses case families against failure classes. Slice-plan mode plans slices before the data exists, with rule-of-three sizing for severe classes. Workflow mode crosses workflow states against case families. Reversibility mode crosses reversibility levels (R0 through R3) against case families.
Why no coverage percentage?
A percentage divides two integers and hides both. Twelve empty cells out of twenty and twelve out of two hundred are different problems; 40% and 6% say less than the numbers they came from.
What does "below the read floor" mean?
A cell with fewer than 16 cases carries a worst-case 95 percent interval wider than ±25 points, which is our default read floor; whether that width is usable depends on the decision the rate feeds.
What is the rule of three?
To claim a severe-failure rate below r with zero observed failures, you need at least n ≥ 3/r cases. The one-sided 95% Clopper-Pearson upper limit at zero failures is 1 − 0.051/n, printed beside the mnemonic so you can see the approximation error.
Why does slice-plan flag sensitive slices?
Declaring a slice means the label has to exist in the data, which for protected-population and vulnerability-signal slices means holding or inferring a sensitive attribute. That is a real tradeoff, not a technicality.