Old harness pass rate
81.00%
INSTRUMENT | Eval statistics
3 cited sources
Separate what the harness moved from what your system moved. Feed the overlap bridge and your eval results to see the harness offset, the system-attributable change and a restated baseline.
Upload an overlap bridge and your baseline and post-migration results. The page separates the harness offset from the system-attributable change and puts an interval on each; with the bridge alone it returns the offset and a restated baseline. The live guide covers the shutdown and the alternatives; this tool takes your files and quantifies the difference. For the exact paired test, see the McNemar calculator.
Showing your last valid result. Update the inputs above to recompute.
System-attributable change (pts)
+1.00
Observed change +11.00 pts, harness offset +10.00 pts, over 1,000 overlap items.
This interval is approximate because the overlap and observed datasets share items.
Old harness pass rate
81.00%
New harness pass rate
91.00%
Every delta is new minus old. The system-attributable change carries an interval and no verdict.
| Stage | Delta (pts) | Interval | Reads as |
|---|---|---|---|
| Harness offset | +10.00 | [+0.0797, +0.1213] | excludes zero |
| Observed change | +11.00 | [+0.0804, +0.1397] | excludes zero |
| System-attributable change | +1.00 | [-0.0264, +0.0460] |
Restated baseline
91.00%
Baseline as recorded
81.00%
The restated baseline falls outside 0 to 100 percent. The arithmetic value is shown as it is.
The observed change is +11.00 points (interval excludes zero). The harness offset accounts for +10.00 points (interval excludes zero). The remainder attributable to the system is +1.00 points, and the interval on that difference is approximate.
Add the dataset named above to complete the attribution.
The link carries the settings only. The eval-set file is not included.
The page runs three stages. First it measures the harness offset from the overlap run: both harnesses scored the same frozen items on the same system, so the only thing that moved between the two columns is the instrument. The overlap 2x2 is old-harness-major, the old harness indexing the rows and the new one the columns, and the interval is the Newcombe paired construction.
Second, it measures the observed change between the baseline and the post-migration results. When both datasets share an item set the page uses a second paired 2x2 on the same pattern. When they do not, it uses the two summary rates and the Newcombe independent construction.
Third, it subtracts the harness offset from the observed change. The system-attributable change is the part of the observed movement that the instrument did not cause.
Every delta is new minus old. A positive number means the rate went up after the change. The harness offset, the observed change and the system-attributable change all follow the same convention.
The interval on the system-attributable change is approximate. The two estimates share items, and the tool does not model that overlap, so the interval can read wider or narrower than the true coverage.
The interval on the system-attributable change is computed by the method of variance estimates recovery (Zou and Donner, 2008). That method assumes the two estimates are independent. The overlap run and the post-migration run can share items, so the true variance carries a covariance term this page does not model. The interval is reported as approximate rather than omitted, and it carries no verdict.
The interval on that difference is approximate. The harness-offset estimate and the observed-change estimate share items, so they are not independent, and the interval construction does not model the covariance between them. A verdict built on an interval that can read too narrow would overstate what the data support.
A flat table with one row per item, an item-id column and one verdict column per harness. CSV and TSV both work. Verdict tokens are case-insensitive: true, 1, yes, y, t for pass and false, 0, no, n, f for fail.
Use it when the baseline and the post-migration run scored the same items and you can match them row by row. The paired interval is narrower because it accounts for the correlation between items that appear in both runs. If the two runs cover different case sets, use the unpaired path with the summary counts.
No. The arithmetic value is shown as it is. A value above 100% or below 0% means the offset is large enough that adding it to the recorded baseline pushes the number past the boundary. It signals that something unexpected happened in the overlap, not a calculation fault.