Criteria
3
INSTRUMENT | Reliability testing
State each pass/fail criterion, the risk axes it covers, the reasoning behind the threshold and who set it. Get the sample size each one needs and export the rules for the release gate designer.
Write each acceptance criterion with the risk it covers (reversibility and blast radius), why the number is that number, who set it, and, for a rate bound, the sample size it needs. When the threshold rests on cost arithmetic, derive the tolerable failure rate first. Built on our statistical significance and judge threshold research. Pairs with the release gate designer. Tighten a gate where the action risk matrix shows wider reach, and split the allowed failures from the SLO error budget.
UNRESOLVED
1
1 unresolved item across 3 criteria.
Criteria
3
Gate rules
3
| Criterion | Form | Reversibility | Blast radius | Gate rule | On fail |
|---|---|---|---|---|---|
| Severe failure rate under 0.1 percent | Bounded tolerance (rate upper bound) | R3: Point of no return. Nothing undoes it. | B3: Outside the organization. Customers, partners, the public, regulated populations. | evidence | hold |
| Core pass rate above 92 percent | Threshold (mandatory floor) | R1: The actor can undo it. Version history, a restorable delete, an idempotent re-apply. | B1: One team, one internal system, one department process. | floor | hold |
| Human review of every irreversible case | Review requirement (evidence) | R3: Point of no return. Nothing undoes it. | B3: Outside the organization. Customers, partners, the public, regulated populations. | evidence | hold |
This tool sizes zero-failure bounds with the one-sided 95 percent Clopper-Pearson bound: the smallest n such that the one-sided upper limit at zero failures is at most the target rate r. This matches the 3/n mnemonic. The mnemonic is slightly conservative (it gives a few more samples than the exact n). An observed rate is reported with the two-sided 95% Clopper-Pearson or Wilson interval.
Each of the ten criterion forms maps to one of the five rule kinds the release gate designer reads (must-pass, floor, regression, budget, evidence) with a fixed on-fail action and a set of fields the gate rule reads. The two risk axes (reversibility and blast radius) are never multiplied or averaged.
| Criterion form | Rule kind | Mandatory | On fail | Fields read |
|---|---|---|---|---|
| Blocker (zero tolerance, rollback) | must-pass | Yes | rollback | none |
| Must-pass (zero tolerance, hold) | must-pass | Yes | hold | none |
| Threshold (mandatory floor) | floor | Yes | hold | direction, threshold |
| Target (advisory) | floor | No | hold | direction, threshold |
| Slice minimum | floor | Yes | hold | direction, threshold, suite |
| Guardrail (budget cap) | budget | Yes | hold | threshold |
| Bounded tolerance (rate upper bound) | evidence | Yes | hold | artifact |
| Statistical condition (regression test) | regression | Yes | hold | tolerance, alpha, baseline |
| Review requirement (evidence) | evidence | Yes | hold | artifact |
| Rollback trigger | must-pass | Yes | rollback | none |
The release gate designer runs the rule against numbers. This tool captures why the rule is set to that number, who decided it, and how many cases the claim needs. A threshold without a stated reason is a threshold nobody can argue with or update.
It lands in the unresolved list at the top. The tool still exports the rule, but the policy document records the gap.
The JSON export writes a document with docType release-gate-spec
and a data payload shaped for the release gate designer's input. Save the
file, open the
release gate designer, and
import it there. The import message is the only record that the thresholds
came from this policy; review each one before locking it in.
The gate cannot evaluate a rate bound directly; it reads the artifact line
and holds until the sample-size evidence is named. The exported rule uses
kind evidence and writes the target rate, required n, and
bound convention into the artifact field.
Floor criteria (threshold, target, slice-minimum) are compared as point estimates by the gate. The confidence interval must be checked by hand. If the interval matters, use a bounded-tolerance or statistical-condition form instead.