LatentEval

INSTRUMENT | Eval statistics

Evaluation Coverage Matrix: Which Cells Have No Cases

Put what you test on one axis and what you test it across on the other. The grid shows which cells have cases, which are empty, which are too small to carry an interval, and what each gap costs you.

Declare what you are testing on one axis and what you are testing it across on another, then see which cells are empty, which are too small to read, and what an empty cell means you do not know. To measure an existing dataset's actual distribution, use the slice and class balance analyzer. Grounded in the testing checklist and the rerun evidence. The risk-weight and persona axes are not built yet. Risk weights come from the action risk matrix and the cost a miss carries, computed in the acceptable error rate calculator.

Row labels
LabelIf empty, you would not know... Actions
Column labels
Label Actions
Cell counts

After adding rows and columns above, enter the count for each cell. Leave a cell blank if no cases exist for that crossing.

Slice planning

The set does not exist yet. This mode plans slices before the data is built.

LevelThresholdWhat it means
TaggedAny countThe slice label exists in the data.
Readable≥ 16 (site default)The worst-case interval is within ±25 points.
Claimable≥ 30A rate claim is supportable (site default, not a derived threshold).
Powered for a named dropSee golden set size plannerThe count the size planner returns for your target rate and effect.

Rule-of-three sizing

To claim a severe-failure rate below r with zero observed failures, you need at least n ≥ 3/r cases. This is the one-sided 95% Clopper-Pearson upper limit at zero failures: 1 − 0.051/n.

Target raten neededExact boundMnemonic (3/n)
1%3000.99%1.00%
0.1%30000.0998%0.100%

Planned slices

Slice familyLabelShare Actions

2 of 9 cells are empty.

Wrong answerRefusal or timeoutSafety violation
Happy path 120 ±8.8% (max) Claimable 40 ±14.8% (max) Claimable 0 Empty Whether the system works under normal conditions at all
Edge input 25 Below claim floor 8 Below read floor 3 Below read floor
Adversarial 0 Empty Whether the system resists deliberately hostile input 15 Below read floor 30 ±16.8% (max) Claimable
Export

A count of empty cells and their positions. No coverage percentage, no grid score: two integers, read as two integers.

cell status = the count against the read floor, the claim floor and the interval it would carryHow?

How this is calculated

Each cell in the grid gets exactly one status. A cell with zero cases is empty. A cell with 1 to 15 cases is below the read floor: too few to carry even a wide confidence interval (16 is the smallest n whose Wald 95% half-width at p = 0.5 is within ±25 points (n ≥ 4z²); cells print the Wilson half-width, which is narrower). A cell with 16 to 29 cases carries a readable interval but a wide one. Whether it supports your claim depends on the rate you want to bound: see slice planning, where n is derived from the target rate rather than a fixed number. Below about 30 cases a rate claim carries an interval wider than most decisions can use; 30 is our default, not a derived threshold. Size the slice from the claim you need with the golden set size planner. A cell with 30 or more cases is claimable, and the widest Wilson half-width the cell could carry (max, computed at p = 0.5) is printed beside the count; a rate away from 0.5 gives a narrower interval.

In slice-plan mode, the rule-of-three sizing shows how many cases you need to claim a rate below a target with zero observed failures. This is the one-sided 95% Clopper-Pearson upper limit at zero failures: n ≥ 3/r (the mnemonic) beside the exact bound 1 − 0.051/n. The acceptance criteria builder sizes the same claim with the two-sided bound, which returns a larger n for the same r (r = 0.001: 3000 here, 3688 there; r = 0.01: 300 here, 368 there). Intersection sparsity multiplies the shares of two slice families to show the expected count at the crossing, which is where the count is smallest and the failure usually is. The product assumes the two families are independent; a correlated pair will be sparser or denser than the product, so treat the figure as a planning floor, not a forecast.

Formula: cell status = the count against the read floor, the claim floor and the interval it would carry

Questions

Questions

What do the four modes cover?

Scenario mode crosses case families against failure classes. Slice-plan mode plans slices before the data exists, with rule-of-three sizing for severe classes. Workflow mode crosses workflow states against case families. Reversibility mode crosses reversibility levels (R0 through R3) against case families.

Why no coverage percentage?

A percentage divides two integers and hides both. Twelve empty cells out of twenty and twelve out of two hundred are different problems; 40% and 6% say less than the numbers they came from.

What does "below the read floor" mean?

A cell with fewer than 16 cases carries a worst-case 95 percent interval wider than ±25 points, which is our default read floor; whether that width is usable depends on the decision the rate feeds.

What is the rule of three?

To claim a severe-failure rate below r with zero observed failures, you need at least n ≥ 3/r cases. The one-sided 95% Clopper-Pearson upper limit at zero failures is 1 − 0.051/n, printed beside the mnemonic so you can see the approximation error.

Why does slice-plan flag sensitive slices?

Declaring a slice means the label has to exist in the data, which for protected-population and vulnerability-signal slices means holding or inferring a sensitive attribute. That is a real tradeoff, not a technicality.