LatentEval

INSTRUMENT | eval

Judge Agreement Tracker: Log Kappa Over Time

1 cited source

Keep a dated log of one judge's agreement with human labels: the value, the prompt version and set behind it, the published band it falls in, the change since the last run, and the next check date.

Keep one judge's agreement figure beside the prompt version and the sample that produced it. Renaming a judge starts a new series. Built on our eval statistics research and on testing whether a judge is reliable.

Bring a number from the agreement calculator, read it against published bands, and pin the judge with a version card. For confidence calibration use the calibration calculator. Entries live in this browser only and can be lost. The Markdown export reads without this site.

The run

Check this value.

A published table is written for one of these, not both.

Check this value.

The judge

The log keeps one series per name. Renaming a judge starts a new one.

Check this value.

Check this value.

Check this value.

Check this value.

Judge card reference, optional

Paste the card reference from the validation report builder. It fills the model, the temperature and the fingerprint, and pins this row to one exact judge.

Check this value.

The set this run scored

A set name, a version, or the fingerprint your golden-set tracker prints.

Check this value.

Check this value.

Check this value.

Check this value.

Optional. The published floor is stated per code, so this is what makes it readable.

Check this value.

The figure

Check this value.

This page never recomputes it. Paste what your calculator gave you.

Check this value.

Stored word for word and never recomputed.

Check this value.

Whose published table this row is interpreted with. Its scope decides whether it applies.

Check this value.

What happens next

Each option is one source's published cadence, not a house recommendation. The cadence table below names the source of each and the date it was read.

Check this value.

Notes

Check this value.

Check this value.

Export this log

Importing replaces everything in this browser with the file's contents. You are asked to confirm, with both counts, and offered a backup first.

delta = value(this run) - value(previous run of the same judge, statistic and pairing)How?

How this is calculated

What counts as the same series. A judge name, an agreement statistic and what was compared. Change any of those and you are looking at a different thing, so the log starts again. The prompt version, the calibration set, the rubric version and the sampling frame are deliberately outside that key: watching a figure move across prompt versions is the whole point of keeping a log, and the published practice draws a fresh set every couple of months, so keying on the set would erase every comparison. Those last three are carried as flags instead. The set reference is free text and this page never checks it, so write the same identifier you used where the set itself is versioned. A dated changelog of what the labeled examples were on the day you froze them is what turns a flagged delta into an answer about which of the two things moved.

The change, and when it is not a like-for-like change. The delta is this run's value minus the previous run's, computed unrounded and shown at three decimals. Where the set, the rubric or the sampling frame moved between the two runs, the delta is flagged and the page says which one moved, because the figure may have moved for that reason rather than because the judge did.

The decline streak. Consecutive falls, counted backwards from the latest run and stopping at the first step that is not a fall. A flat step stops it. The page prints the count and passes no verdict of its own: the one rolling-drop rule in the cadence table reads "2 point rolling drop" and its source does not settle whether that means two hundredths of kappa or two points of raw agreement, so there is no threshold here to cite.

The next check. Each schedule option is one source's own published cadence. A monthly cadence adds a calendar month and clamps to the last day of the target month, so a run on the 31st of January falls due on the 28th of February, or the 29th in a leap year. A schedule printed as a range resolves to its earliest day. A source that schedules on an event rather than on the calendar computes no date at all.

The interpretation column. Saving a row stores the whole band reading: the verdict, the printed range, the source, the scope that source claims, and the date a person read it. The row keeps printing what the table said on that date. Where the live table now disagrees, both readings and both dates are shown side by side rather than the newer one quietly replacing the older. A table that was not written for your statistic refuses rather than guessing, and names the sources that did write one.

The sample size. This page passes no verdict on it. The nearest published floor is stated per error code and this form records a total, so a total cannot be tested against it. Fill in the number of error codes and the page divides, which is the only honest way to put the two beside each other, and even then it is an average: a thin code hides inside a healthy one.

Three sources, side by side. The spread is the finding. A vendor source is named with the date a person read it and carries no link.
What the source schedules Set size How often Source Read on
Initial calibration, dual labeled 200 once Future AGI 2026-08-21
Drift check, fresh set 50 monthly Future AGI 2026-08-21
Alert on rolling kappa drop none stated 2 point rolling drop, per model version change Future AGI 2026-08-21
Annotations per error code 15 to 20 per calibration round AWS sample-GEDD 2026-08-21
Recalibration fresh annotation set quarterly, every 60 to 90 days AWS sample-GEDD 2026-08-21
Production trace sample 100 to 300 per calibration loop Future AGI 2026-08-21
Canary eval against fixed ground truth none stated weekly Galileo 2026-08-21
Human spot check, stratified none stated monthly Galileo 2026-08-21
Full calibration cycle with SME annotation none stated quarterly, or on a red-flag signal Galileo 2026-08-21

Formula: delta = value(this run) - value(previous run of the same judge, statistic and pairing)

What the worksheet contains

The Markdown export is an annotator calibration worksheet. It opens in any plain editor and reads as a document: the judge and the prompt version behind the figure, the set and the sampling that produced it, the result with its interval as you recorded it, the band it fell in with the source and the date that source was read, the next check date, and a table of every run in the log with its change and its band verdict. It points at nothing that only resolves on this site, so it survives being pasted into a repository or a review doc.

The JSON export is the same log in the shape this page reads back. There is one JSON shape, so the file you export is the file the import control accepts. The CSV is the log as a spreadsheet, one row per run, with every column the table above shows.

This log holds one judge's agreement figures and the identity behind them. The runs those figures were computed over, with the dataset, the pass rate and the plan frozen before anyone looked, belong in the register that keeps one row per eval run, which exports the same way, so the two files sit side by side in a repository without either one needing this site.

Questions

Questions

Is this the same calibration as the judge calibration calculator?

No, and the word carries two senses on this site. That calculator scores whether a judge's stated confidence tracks how often it is actually right, with a Brier score and expected calibration error. This page keeps the record of how well a judge agrees with human labels, measured again over time. If you have confidence scores rather than an agreement figure, the calibration calculator is the page you want.

Why will this page not compute the coefficient for me?

Because computing it well is a page of its own. Bring the figure from the agreement calculator, which handles three or more raters, missing cells and interval estimates, or from the two-rater kappa calculator. This page never recomputes a number you bring, the interval included: what you recorded is what it keeps.

I renamed a judge and the log started again. Why?

The page cannot tell a rename from a new judge, and guessing wrong is worse in both directions: treating two judges as one invents a drift that never happened, and treating one judge as two hides a real one. So a name is an identity. If you rename, export first, change the name in the file, and import it back.

Which published table should I read my figure against?

The one whose scope matches what you did. The tables in the select disagree with each other, and they were written for different things: some for two human annotators, some for a judge against one human, some for one statistic only. Each stores its scope with your row, so the row can be read honestly later. A table that was not written for your statistic refuses rather than returning a verdict, which is the point of asking. There is more on where the numbers come from in our write-up on kappa thresholds.

What happens when I hit the limit?

The log holds 500 runs. At the limit the page refuses the add and offers you the export and the clear control rather than silently dropping the oldest row. Browser storage is a few megabytes shared across every tool on this site, and a browser evicts an origin all at once, so a log with no ceiling would take its neighbors with it.

Does anything I type leave the browser?

No. There is no server behind this page, no account and no upload. Judge names, prompt versions and set references are often the parts you can least afford to send anywhere, which is why the record lives where it does. The page does count anonymous usage: which tool was opened, which export button was pressed, and nothing from any field.

Sources

  1. MDN, Storage quotas and eviction criteriaMDN Web Docs, Mozilla Retrieved