LatentEval

INSTRUMENT | Reliability testing

Acceptance Criteria Builder: Set the Bar Before the Run

State each pass/fail criterion, the risk axes it covers, the reasoning behind the threshold and who set it. Get the sample size each one needs and export the rules for the release gate designer.

Write each acceptance criterion with the risk it covers (reversibility and blast radius), why the number is that number, who set it, and, for a rate bound, the sample size it needs. When the threshold rests on cost arithmetic, derive the tolerable failure rate first. Built on our statistical significance and judge threshold research. Pairs with the release gate designer. Tighten a gate where the action risk matrix shows wider reach, and split the allowed failures from the SLO error budget.

UNRESOLVED

1

1 unresolved item across 3 criteria.

Criteria

3

Gate rules

3

Policy

Check this value.

Criterion 1

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Required sample size: 2995 cases (one-sided 95% Clopper-Pearson upper bound)

Rule-of-three mnemonic (one-sided 95% bound): 3/0.001 = 3000

One-sided 95% upper bound at n = 2995: 0.1%

The 3/n mnemonic uses the same one-sided 95% convention as this tool. At the mnemonic's n = 3000 the bound is 0.0998%, at or below the 0.1% target.

Criterion 2

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

This criterion gates on a point estimate. The gate cannot evaluate a rate bound directly. Consider using a bounded-tolerance or statistical-condition form.

Criterion 3

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Check this value.

Criteria mapping

Criterion Form Reversibility Blast radius Gate rule On fail
Severe failure rate under 0.1 percent Bounded tolerance (rate upper bound) R3: Point of no return. Nothing undoes it. B3: Outside the organization. Customers, partners, the public, regulated populations. evidence hold
Core pass rate above 92 percent Threshold (mandatory floor) R1: The actor can undo it. Version history, a restorable delete, an idempotent re-apply. B1: One team, one internal system, one department process. floor hold
Human review of every irreversible case Review requirement (evidence) R3: Point of no return. Nothing undoes it. B3: Outside the organization. Customers, partners, the public, regulated populations. evidence hold
Export the policy
Export rules as release-gate-spec
How this is calculated

Sample-size calculation

This tool sizes zero-failure bounds with the one-sided 95 percent Clopper-Pearson bound: the smallest n such that the one-sided upper limit at zero failures is at most the target rate r. This matches the 3/n mnemonic. The mnemonic is slightly conservative (it gives a few more samples than the exact n). An observed rate is reported with the two-sided 95% Clopper-Pearson or Wilson interval.

Form-to-rule mapping

Each of the ten criterion forms maps to one of the five rule kinds the release gate designer reads (must-pass, floor, regression, budget, evidence) with a fixed on-fail action and a set of fields the gate rule reads. The two risk axes (reversibility and blast radius) are never multiplied or averaged.

Criterion form Rule kind Mandatory On fail Fields read
Blocker (zero tolerance, rollback) must-pass Yes rollback none
Must-pass (zero tolerance, hold) must-pass Yes hold none
Threshold (mandatory floor) floor Yes hold direction, threshold
Target (advisory) floor No hold direction, threshold
Slice minimum floor Yes hold direction, threshold, suite
Guardrail (budget cap) budget Yes hold threshold
Bounded tolerance (rate upper bound) evidence Yes hold artifact
Statistical condition (regression test) regression Yes hold tolerance, alpha, baseline
Review requirement (evidence) evidence Yes hold artifact
Rollback trigger must-pass Yes rollback none

Questions

Why does this tool exist alongside the release gate?

The release gate designer runs the rule against numbers. This tool captures why the rule is set to that number, who decided it, and how many cases the claim needs. A threshold without a stated reason is a threshold nobody can argue with or update.

What happens when a criterion has no sample size?

It lands in the unresolved list at the top. The tool still exports the rule, but the policy document records the gap.

How does the release-gate-spec export work?

The JSON export writes a document with docType release-gate-spec and a data payload shaped for the release gate designer's input. Save the file, open the release gate designer, and import it there. The import message is the only record that the thresholds came from this policy; review each one before locking it in.

How does the gate handle a bounded-tolerance criterion?

The gate cannot evaluate a rate bound directly; it reads the artifact line and holds until the sample-size evidence is named. The exported rule uses kind evidence and writes the target rate, required n, and bound convention into the artifact field.

What about floor and threshold criteria?

Floor criteria (threshold, target, slice-minimum) are compared as point estimates by the gate. The confidence interval must be checked by hand. If the interval matters, use a bounded-tolerance or statistical-condition form instead.