LatentEval

INSTRUMENT | Eval statistics

Harness Migration Score Continuity Checker

3 cited sources

Separate what the harness moved from what your system moved. Feed the overlap bridge and your eval results to see the harness offset, the system-attributable change and a restated baseline.

Upload an overlap bridge and your baseline and post-migration results. The page separates the harness offset from the system-attributable change and puts an interval on each; with the bridge alone it returns the offset and a restated baseline. The live guide covers the shutdown and the alternatives; this tool takes your files and quantifies the difference. For the exact paired test, see the McNemar calculator.

What you have
Overlap run

Both harnesses scored the same frozen items on the same day, using the same model, prompts and scaffold the baseline number came from. A bridge measured on the post-migration system does not restate a baseline. The file is a flat table with one row per item: an item id column and one verdict column per harness.

Overlap source
Overlap counts

Items the old harness passed and the new harness passed.

Check this value.

Items the old harness passed and the new harness failed.

Check this value.

Items the old harness failed and the new harness passed.

Check this value.

Items the old harness failed and the new harness failed.

Check this value.

Observed change
Baseline and post-migration runs
Summary counts

Items that passed after the port.

Check this value.

Items scored after the port.

Check this value.

Items that passed before the port.

Check this value.

Items scored before the port.

Check this value.

Check this value.

To export your data into the new harness's format, see the schema validator.

System-attributable change (pts)

+1.00

Observed change +11.00 pts, harness offset +10.00 pts, over 1,000 overlap items.

This interval is approximate because the overlap and observed datasets share items.

Old harness pass rate

81.00%

New harness pass rate

91.00%

Every delta is new minus old. The system-attributable change carries an interval and no verdict.

StageDelta (pts)IntervalReads as
Harness offset +10.00 [+0.0797, +0.1213] excludes zero
Observed change +11.00 [+0.0804, +0.1397] excludes zero
System-attributable change +1.00 [-0.0264, +0.0460]

Restated baseline

91.00%

Baseline as recorded

81.00%

The observed change is +11.00 points (interval excludes zero). The harness offset accounts for +10.00 points (interval excludes zero). The remainder attributable to the system is +1.00 points, and the interval on that difference is approximate.

Export

The link carries the settings only. The eval-set file is not included.

How this is calculated

The page runs three stages. First it measures the harness offset from the overlap run: both harnesses scored the same frozen items on the same system, so the only thing that moved between the two columns is the instrument. The overlap 2x2 is old-harness-major, the old harness indexing the rows and the new one the columns, and the interval is the Newcombe paired construction.

Second, it measures the observed change between the baseline and the post-migration results. When both datasets share an item set the page uses a second paired 2x2 on the same pattern. When they do not, it uses the two summary rates and the Newcombe independent construction.

Third, it subtracts the harness offset from the observed change. The system-attributable change is the part of the observed movement that the instrument did not cause.

Every delta is new minus old. A positive number means the rate went up after the change. The harness offset, the observed change and the system-attributable change all follow the same convention.

The interval on the system-attributable change is approximate. The two estimates share items, and the tool does not model that overlap, so the interval can read wider or narrower than the true coverage.

The interval on the system-attributable change is computed by the method of variance estimates recovery (Zou and Donner, 2008). That method assumes the two estimates are independent. The overlap run and the post-migration run can share items, so the true variance carries a covariance term this page does not model. The interval is reported as approximate rather than omitted, and it carries no verdict.

Questions

Why does the system-attributable change carry no verdict?

The interval on that difference is approximate. The harness-offset estimate and the observed-change estimate share items, so they are not independent, and the interval construction does not model the covariance between them. A verdict built on an interval that can read too narrow would overstate what the data support.

What file format does the overlap run need?

A flat table with one row per item, an item-id column and one verdict column per harness. CSV and TSV both work. Verdict tokens are case-insensitive: true, 1, yes, y, t for pass and false, 0, no, n, f for fail.

When should I use the paired observed path instead of the unpaired one?

Use it when the baseline and the post-migration run scored the same items and you can match them row by row. The paired interval is narrower because it accounts for the correlation between items that appear in both runs. If the two runs cover different case sets, use the unpaired path with the summary counts.

The restated baseline is outside 0 to 100 percent. Is that an error?

No. The arithmetic value is shown as it is. A value above 100% or below 0% means the offset is large enough that adding it to the recorded baseline pushes the number past the boundary. It signals that something unexpected happened in the overlap, not a calculation fault.

Sources

  1. Construction of confidence limits about effect measures: a general approachStatistics in Medicine 27(10), 1693-1702 (Zou and Donner, 2008) Retrieved
  2. Interval estimation for the difference between independent proportions: comparison of eleven methodsStatistics in Medicine 17(8), 873-890 (Newcombe, 1998) Retrieved
  3. Improved confidence intervals for the difference between binomial proportions based on paired dataStatistics in Medicine 17(22), 2635-2650 (Newcombe, 1998) Retrieved