Change since the previous run
none yet
INSTRUMENT | eval
1 cited source
Keep a dated log of one judge's agreement with human labels: the value, the prompt version and set behind it, the published band it falls in, the change since the last run, and the next check date.
Keep one judge's agreement figure beside the prompt version and the sample that produced it. Renaming a judge starts a new series. Built on our eval statistics research and on testing whether a judge is reliable.
Bring a number from the agreement calculator, read it against published bands, and pin the judge with a version card. For confidence calibration use the calibration calculator. Entries live in this browser only and can be lost. The Markdown export reads without this site.
Showing your last valid result. Update the inputs above to recompute.
Latest standing
0.000
No reading yet
Change since the previous run
none yet
Next check due
none set
Every run recorded in this browser
| Run date | Judge and version | Set | Statistic | Value | Change | Comparable | Interpretation | Row actions |
|---|
Importing replaces everything in this browser with the file's contents. You are asked to confirm, with both counts, and offered a backup first.
delta = value(this run) - value(previous run of the same judge, statistic and pairing)How?
What counts as the same series. A judge name, an agreement statistic and what was compared. Change any of those and you are looking at a different thing, so the log starts again. The prompt version, the calibration set, the rubric version and the sampling frame are deliberately outside that key: watching a figure move across prompt versions is the whole point of keeping a log, and the published practice draws a fresh set every couple of months, so keying on the set would erase every comparison. Those last three are carried as flags instead. The set reference is free text and this page never checks it, so write the same identifier you used where the set itself is versioned. A dated changelog of what the labeled examples were on the day you froze them is what turns a flagged delta into an answer about which of the two things moved.
The change, and when it is not a like-for-like change. The delta is this run's value minus the previous run's, computed unrounded and shown at three decimals. Where the set, the rubric or the sampling frame moved between the two runs, the delta is flagged and the page says which one moved, because the figure may have moved for that reason rather than because the judge did.
The decline streak. Consecutive falls, counted backwards from the latest run and stopping at the first step that is not a fall. A flat step stops it. The page prints the count and passes no verdict of its own: the one rolling-drop rule in the cadence table reads "2 point rolling drop" and its source does not settle whether that means two hundredths of kappa or two points of raw agreement, so there is no threshold here to cite.
The next check. Each schedule option is one source's own published cadence. A monthly cadence adds a calendar month and clamps to the last day of the target month, so a run on the 31st of January falls due on the 28th of February, or the 29th in a leap year. A schedule printed as a range resolves to its earliest day. A source that schedules on an event rather than on the calendar computes no date at all.
The interpretation column. Saving a row stores the whole band reading: the verdict, the printed range, the source, the scope that source claims, and the date a person read it. The row keeps printing what the table said on that date. Where the live table now disagrees, both readings and both dates are shown side by side rather than the newer one quietly replacing the older. A table that was not written for your statistic refuses rather than guessing, and names the sources that did write one.
The sample size. This page passes no verdict on it. The nearest published floor is stated per error code and this form records a total, so a total cannot be tested against it. Fill in the number of error codes and the page divides, which is the only honest way to put the two beside each other, and even then it is an average: a thin code hides inside a healthy one.
| What the source schedules | Set size | How often | Source | Read on |
|---|---|---|---|---|
| Initial calibration, dual labeled | 200 | once | Future AGI | 2026-08-21 |
| Drift check, fresh set | 50 | monthly | Future AGI | 2026-08-21 |
| Alert on rolling kappa drop | none stated | 2 point rolling drop, per model version change | Future AGI | 2026-08-21 |
| Annotations per error code | 15 to 20 | per calibration round | AWS sample-GEDD | 2026-08-21 |
| Recalibration | fresh annotation set | quarterly, every 60 to 90 days | AWS sample-GEDD | 2026-08-21 |
| Production trace sample | 100 to 300 | per calibration loop | Future AGI | 2026-08-21 |
| Canary eval against fixed ground truth | none stated | weekly | Galileo | 2026-08-21 |
| Human spot check, stratified | none stated | monthly | Galileo | 2026-08-21 |
| Full calibration cycle with SME annotation | none stated | quarterly, or on a red-flag signal | Galileo | 2026-08-21 |
Formula: delta = value(this run) - value(previous run of the same judge, statistic and pairing)
The Markdown export is an annotator calibration worksheet. It opens in any plain editor and reads as a document: the judge and the prompt version behind the figure, the set and the sampling that produced it, the result with its interval as you recorded it, the band it fell in with the source and the date that source was read, the next check date, and a table of every run in the log with its change and its band verdict. It points at nothing that only resolves on this site, so it survives being pasted into a repository or a review doc.
The JSON export is the same log in the shape this page reads back. There is one JSON shape, so the file you export is the file the import control accepts. The CSV is the log as a spreadsheet, one row per run, with every column the table above shows.
This log holds one judge's agreement figures and the identity behind them. The runs those figures were computed over, with the dataset, the pass rate and the plan frozen before anyone looked, belong in the register that keeps one row per eval run, which exports the same way, so the two files sit side by side in a repository without either one needing this site.
No, and the word carries two senses on this site. That calculator scores whether a judge's stated confidence tracks how often it is actually right, with a Brier score and expected calibration error. This page keeps the record of how well a judge agrees with human labels, measured again over time. If you have confidence scores rather than an agreement figure, the calibration calculator is the page you want.
Because computing it well is a page of its own. Bring the figure from the agreement calculator, which handles three or more raters, missing cells and interval estimates, or from the two-rater kappa calculator. This page never recomputes a number you bring, the interval included: what you recorded is what it keeps.
The page cannot tell a rename from a new judge, and guessing wrong is worse in both directions: treating two judges as one invents a drift that never happened, and treating one judge as two hides a real one. So a name is an identity. If you rename, export first, change the name in the file, and import it back.
The one whose scope matches what you did. The tables in the select disagree with each other, and they were written for different things: some for two human annotators, some for a judge against one human, some for one statistic only. Each stores its scope with your row, so the row can be read honestly later. A table that was not written for your statistic refuses rather than returning a verdict, which is the point of asking. There is more on where the numbers come from in our write-up on kappa thresholds.
The log holds 500 runs. At the limit the page refuses the add and offers you the export and the clear control rather than silently dropping the oldest row. Browser storage is a few megabytes shared across every tool on this site, and a browser evicts an origin all at once, so a log with no ceiling would take its neighbors with it.
No. There is no server behind this page, no account and no upload. Judge names, prompt versions and set references are often the parts you can least afford to send anywhere, which is why the record lives where it does. The page does count anonymous usage: which tool was opened, which export button was pressed, and nothing from any field.