Dataset version
not named
INSTRUMENT | Reliability testing
3 cited sources
Write the rule that turns your eval figures into PROMOTE, HOLD or ROLLBACK, get the verdict plus the rule that produced it, and keep a dated decision log and memo in your browser.
A release gate is a written rule that turns eval figures into one of three decisions. Write the rule here and keep the decision. This page runs no suite.
You get a verdict, the rule behind it, and a line to paste into the release channel. One number cannot say which rule failed, so there is none.
Built on our work on whether a difference is real and what a pass rate hides. Pairs with the change-impact worksheet.
Showing your last valid result. Update the inputs above to recompute.
Mandatory rules not passing
1
HOLD. Task success regression did not evaluate.
HOLD: evidence missing for Task success regression
1 of 5 mandatory rules does not pass: 1 did not evaluate.
Every rule, as written, with what it read
| Rule | Suite | Rule as written | Observed | Result | On failure |
|---|---|---|---|---|---|
| Destructive-action safety cases | safety-critical | no case in this suite may fail | 0 failing | pass | |
| Task success floor | core-task | observed at or above 0.8 | 0.85 | pass | |
| Task success regression | core-task | a regression rule with its direction or unit unset | 0.842 | not evaluated | hold |
| Remaining error budget | slo | remaining at or above 0 | 1 | pass | |
| Rollback or compensation path | release-ops | an artifact is named | revert to previous tag | pass |
Notes
Contradictions
Dataset version
not named
Frozen on
not named
Nothing imported yet. A regression rule still evaluates without this; the memo then records the baseline set as unnamed.
A frozen value that arrives here is cited, never compared. Freezing one is the version tracker's job, and a figure moves into a rule above only because you typed or pasted it.
Decisions stay in this browser and go nowhere else. Up to 200 of them; after that, export and clear.
Decision log
| Date | Candidate | Baseline | Verdict | Deciding rule | Decided by | Override | Dataset version |
|---|---|---|---|---|---|---|---|
| No decisions saved yet. | |||||||
This deletes every saved decision in this browser. Press clear again to confirm.
Bring a log back
Reads a log you exported from this page and replaces what is here. Export the current one first: replacing is not a merge.
JSON. Drag one here, or use the box below.
Check this value.
verdict = the worst on-failure state among the rules that failedHow?
Two steps, and neither one combines rules. Each rule is judged on its own terms: a must-pass rule passes when its failing-case count is zero; a floor compares one observed value against one threshold; a regression compares the movement from a baseline against a tolerance, in points or as a percent of that baseline; a budget compares what is left against a line; an evidence rule passes when an artifact is named.
Then the verdict is the worst on-failure state among the rules that failed, and the deciding rule is the first one at that level in your own row order. A rule that could not be evaluated counts as failing and holds. Every comparison is inclusive and carries a small tolerance for floating point, so a deterioration of exactly 0.020 against a tolerance of 0.020 passes rather than failing on a value the machine stores as 0.020000000000000018.
A rule can only soften from roll back to hold when both of two things hold: your p-value is above your alpha, and the movement is inside the noise floor you stated. A p-value on its own is not evidence that nothing changed.
No number here spans more than one rule. Four rules passing out of six is reported as a count of what failed, never as a proportion, because a proportion cannot say which rule it was.
Formula: verdict = the worst on-failure state among the rules that failed
One entry. It is one team's convention on one system, shown so you can see what a published line looks like. Nothing here is filled into a rule above.
Rollback line at 70 percent of the target
As stated: a build is rejected outright when any one dimension falls below 70 percent of its own target, for example below 56 percent where the target is 80 percent.
Claimed for one team’s staging gate on one internally deployed multi-agent assistant, across 38 evaluation runs over four weeks, with five named quality dimensions.
Automated Self-Testing as a Quality Gate (arXiv 2603.15676), read 2026-08-27
One team’s line on its own system. It is not carried over to yours, and this page never fills it into a rule for you.
Because one number cannot tell you which rule failed, and that is the only part of a gate a release channel can act on. A page that averaged six rules into a percentage would let a failed safety case hide behind five passes.
They disagree. A drop from 0.860 to 0.842 is 0.018 points, which passes a 0.020-point tolerance, and 2.09 percent of the baseline, which fails a 2-percent one. Pick the one your team argues in, and write it down.
Because the wrong orientation passes a worse release quietly. On latency, higher is worse; on task success, higher is better. There is no safe default, so the rule holds instead of guessing.
No. It reads figures you bring, from your own runs. There is no repository field, no workflow hook, and nothing to connect.
Into this browser, and nowhere else. Export them before you clear the browser, change machine, or hand the record to anyone.