INSTRUMENT | Reliability testing
Eval Run Register: Preregistration and Run Cards for LLM Evals
3 cited sources
Keep one row per eval run in this browser: freeze the plan before results exist, log what happened, and export a run card or a preregistration record for a PR.
One row per eval run, exported as a run card or preregistration record: our routing refusal tax study, Fable 5 vs Opus 5 vs Opus 4.8, and Fable 5 vs Sol vs Kimi K3 benchmarks are its first entries. Evaluation cards are separate.
A rule fixed after results is not a rule, so the plan freezes first, and a later edit is named field by field. See the four ways an eval number misleads. It runs, sends, and prices nothing.
Start a run record, or load our own study as a filled-in example.
Baseline delta
Verdict: not yet run
Check this value.
Check this value.
Check this value.
Check this value.
Check this value.
Check this value.
Check this value.
Check this value.
Results for IDX
Check this value.
Check this value.
Check this value.
Check this value.
Check this value.
Check this value.
Check this value.
Check this value.
Check this value.
How this is calculated
Verdict. Read the subject variant's reading whose metric matches the primary metric, and compare it to the threshold with the chosen rule (at least, at most, greater than, less than). No rule chosen yields "inconclusive" and the reader sets a verdict by hand with a reason. Numbers compare raw, never rounded first. One threshold on one metric is a reading, not a shipping decision: where several rules have to agree before anything goes out, write them down in the rule set that turns figures into promote, hold or roll back and cite this entry as the evidence behind one of them.
Plan fingerprint. The plan is turned into one deterministic string (sorted fields, collapsed whitespace, variants sorted by their id so reordering them on screen is not a change) and hashed with SHA-256. Freezing stores that hash plus a full copy of the plan. The dataset version in that plan is a string this page takes on trust, so it is worth citing one you can look up later: what actually changed between two versions of an eval set is a separate record, and a plan that names a version nobody kept is a plan nobody can rerun.
Plan drift. Every later save recomputes the fingerprint over the live plan. A mismatch names exactly which fields changed since the freeze, by comparing the live plan to the frozen copy field by field. Once flagged, an entry stays flagged, even if the edit is reverted.
Baseline difference. Subject value minus baseline value on the primary metric, shown exactly as entered. It is arithmetic, not a significance test, and the page never claims otherwise.
Missing-field rate. Over completed entries only, the share missing at least one of seven fields a re-run needs (model id, model version, prompt version, temperature, max tokens, dataset version, and the sample size behind the primary reading), to one decimal place.
Questions
What does the fingerprint prove, and what does it not?
It proves the plan on this device has not changed since the stamp it is compared against. It proves nothing to anyone else about when that stamp was made, because the reader controls both the clock and the file. Third-party proof needs a trusted timestamp from an independent party, a time-stamping authority, per IETF RFC 3161, which this page does not provide, and it names no product for one.
What does the preregistration export contain?
Its JSON data payload is the frozen plan itself, plus the freeze timestamp and the full fingerprint, so anyone can recompute the fingerprint from the document alone with no results section attached. Read more on why a plan written after results has stopped being a plan at our methodology page.
How does a two-variant run round-trip?
Every result attaches to the variant that produced it by a minted id, never by its editable label, so a rename never detaches a result. The register names which variant the verdict is about (the subject) and which is the baseline, and the difference between them carries no significance claim.
A run judged by a model carries a second question this register does not answer: whether the judge still agrees with people. That belongs in a dated log of one judge's agreement figures, which keeps its own series per judge and flags the runs where the prompt, the model or the calibration set moved under it.