Kappa thresholds for LLM judges, and who published each one
Five published kappa bands from four sources, side by side, each with the author who wrote it, the date we read it, and the coefficient it was written for. They are conventions, and they disagree.
Part of LLM-as-a-judge bias, and the tests that catch it
In brief
5 POINTS- No published kappa mapping is a standard. Five of them exist, from four sources, and they disagree on the same number.
- A band written for Cohen’s kappa does not transfer to Gwet’s AC1 or Fleiss’ kappa. Each table below names the coefficient its author wrote it for.
- 0.65 reads substantial on Landis and Koch and moderate on McHugh 2012, which was written to argue that Landis and Koch is too lenient.
- Future AGI publishes no judge-to-human band at or below 0.6, so a value there gets no reading rather than an invented one.
- Recalibration cadences run from weekly to quarterly across three vendors. The spread is the finding.
On this page (5)
There is no standard kappa for an LLM judge. There are conventions, five of them from four sources, and they disagree about what the same number is worth. This page prints all five with the author of each, the date we read it, and the coefficient it was written for, then does the same for how often to check a judge again.
This page rests on three pieces of our own work, and it has to reconcile with two of them. Our checklist for auditing an LLM judge gates a ranking on a Cohen’s kappa of 0.61 and names Landis and Koch as where that figure comes from. The reliability testing checklist for agents restates the same band as a pass line. And our hub for eval statistics carries the prevalence result that the last section here depends on. So two of our own pages publish a pass line at 0.61. The table below is where that number came from, who else disagrees with it, and what it was written to measure.
We publish no house cutoff of our own. Each table below is a convention its author proposed for a particular kind of work, and setting a house number would mean averaging four of them across different coefficients, different pairings and different fields, then calling the average a standard.
The five published mappings
Each row names the source and the day we read it. The last column is the part most readers skip and should not: a band written for Cohen’s kappa between two humans is not a reading for a Gwet’s AC1 between a judge and a human, and three of these five sources say so about themselves.
Published interpretation bands for chance-corrected agreement coefficients
| Range | What the source says |
|---|---|
| Landis and Koch 1977, via AHRQ, read 2026-08-21. Written for Cohen's kappa, Fleiss' kappa, Gwet's AC1, Krippendorff's alpha, either pairing. | |
| < 0 | Poor |
| 0.0 to 0.20 | Slight |
| 0.21 to 0.40 | Fair |
| 0.41 to 0.60 | Moderate |
| 0.61 to 0.80 | Substantial |
| 0.81 to 1.0 | Almost perfect |
| McHugh 2012, Biochemia Medica, read 2026-08-21. Written for Cohen's kappa, human against human. | |
| 0 to .20 | None, 0 to 4 percent of data reliable |
| .21 to .39 | Minimal, 4 to 15 percent |
| .40 to .59 | Weak, 15 to 35 percent |
| .60 to .79 | Moderate, 35 to 63 percent |
| .80 to .90 | Strong, 64 to 81 percent |
| above .90 | Almost perfect, 82 to 100 percent |
| AWS sample-GEDD, read 2026-08-21. Written for Cohen's kappa, judge against human. | |
| < 0.00 | Rubric is inverted, fix immediately |
| 0.00 to 0.20 | Judge is unreliable, do not use |
| 0.21 to 0.40 | Major rubric revision needed |
| 0.41 to 0.60 | Usable with human review on flagged cases |
| 0.61 to 0.79 | Acceptable for low-stakes automation |
| >= 0.80 | Deploy autonomously in CI |
| Future AGI, read 2026-08-21. Written for Cohen's kappa, Fleiss' kappa, human against human. | |
| below 0.4 | The rubric is ambiguous, rewrite it |
| 0.4 to 0.6 | Weak, the rubric is tunable |
| above 0.6 | Acceptable |
| above 0.8 | Strong rubric |
| Future AGI, read 2026-08-21. Written for Cohen's kappa, Fleiss' kappa, judge against human. | |
| above 0.6 | Acceptable for production |
| above 0.8 | Strong |
Two things fall out of reading it as a block. The same number gets different words: 0.65 is substantial on Landis and Koch and moderate on McHugh 2012. And one of the five is not a vocabulary at all. The AWS mapping replaces the adjective with a decision, so its cells say what to do rather than what the number is worth.
You can run the coefficients themselves in the multi-rater agreement calculator, which reads this same table and will only apply a set that names the coefficient it just computed. If you are still choosing which check to run, the judge agreement and bias instruments pair each situation with the tool that answers it. The agreement tracker’s own interpretation column reads its wording from this same set of tables, so a coefficient logged there carries the verdict this page would give it.
Why they disagree
Landis and Koch 1977, via AHRQ. The oldest and the most repeated. It was written for observer agreement on categorical medical data, and the 1977 paper was written for kappa. It is the lineage almost every later table adapts, and it is usually quoted with no source attached at all. We read AC1 and Krippendorff’s alpha on it here because a published annotation study does the same. The 1977 paper does not say you may.
McHugh 2012, Biochemia Medica. Written to argue that the Landis and Koch reading is too lenient for research where a wrong call has consequences. It is the strictest of the five, it carries a percent-of-data-reliable column beside each band, and it prints no band at all below zero. Two raters, Cohen’s kappa, healthcare research.
AWS sample-GEDD. The only set that ties a band to an action. Its scope is narrow and stated: Cohen’s kappa, one LLM judge against one human annotator, per criterion, on grounded agent evals. That narrowness is why it is the most useful of the five inside its scope and the most dangerous outside it. A deploy instruction written for a judge-against-human Cohen’s kappa says nothing about a three-rater Fleiss’ kappa or a Gwet’s AC1.
Future AGI, inter-annotator. Agreement between human annotators on a judge rubric, which is a different question from agreement between the judge and a human. The source says kappa without naming the variant and describes two or three labelers, so we list it for both Cohen’s and Fleiss’ kappa and leave the ambiguity visible.
Future AGI, judge to human. The same publisher’s separate targets for the judge-against-human pairing. It publishes nothing at or below 0.6. A lookup down there returns no band, and this page prints that rather than inventing one, because the gap is a fact about the source.
Two readings of the same numbers
One dataset, two coefficients, two verdicts. An annotation study reports 95.05 percent raw agreement on its labels, with a Cohen’s kappa of 0.495 and a Gwet’s AC1 of 0.945 on that same data. On Landis and Koch the kappa reads moderate and the AC1 reads almost perfect. On the AWS mapping the kappa reads usable with human review on flagged cases, and the AC1 reads nothing: AWS wrote for Cohen’s kappa, so the lookup refuses rather than lending its wording to a coefficient its author never tested. The gap between 0.495 and 0.945 is the label mix, and it is the reason a single band table cannot answer for both figures at once.
A judge that agrees with nothing. A widely cited worked case has 100 outputs, 95 of them safe, and 90 percent raw agreement between a judge and a human. Chance agreement on that label mix is 90.5 percent, so the kappa lands at about negative 0.05. Landis and Koch calls that poor. AWS calls it a rubric that is inverted and should be fixed immediately. McHugh returns nothing, because its scale starts at zero. Three sources, three treatments of one number, and the most actionable of them is the one with the narrowest declared scope.
How often to check again
Agreement is a measurement of a judge at a moment. The prompt changes, the model version changes, the traffic changes, and the number goes stale. Three vendors publish a cadence and none of them agrees with the others, so the honest reading is the range rather than any one row.
Published calibration and drift-check cadences for LLM judges
| What to run | How many | How often | Source | What that source was writing about |
|---|---|---|---|---|
| Initial calibration, dual labeled | 200 | once | Future AGI, read 2026-08-21 | Hybrid judge plus human verification, framed by the source as a 2026 default. |
| Drift check, fresh set | 50 | monthly | Future AGI, read 2026-08-21 | Hybrid judge plus human verification, framed by the source as a 2026 default. |
| Alert on rolling kappa drop | not stated | 2 point rolling drop, per model version change | Future AGI, read 2026-08-21 | Hybrid judge plus human verification, framed by the source as a 2026 default. |
| Annotations per error code | 15 to 20 | per calibration round | AWS sample-GEDD, read 2026-08-21 | Per criterion, on grounded agent evals. |
| Recalibration | fresh annotation set | quarterly, every 60 to 90 days | AWS sample-GEDD, read 2026-08-21 | Per criterion, on grounded agent evals. |
| Production trace sample | 100 to 300 | per calibration loop | Future AGI, read 2026-08-21 | Production traces, with two or three human labelers. |
| Canary eval against fixed ground truth | not stated | weekly | Galileo, read 2026-08-21 | Production judge programs. |
| Human spot check, stratified | not stated | monthly | Galileo, read 2026-08-21 | Production judge programs. |
| Full calibration cycle with SME annotation | not stated | quarterly, or on a red-flag signal | Galileo, read 2026-08-21 | Production judge programs. |
The spread runs from a weekly canary to a quarterly full cycle, and the rows agree on less than that range suggests. Five of the nine state a sample size, and four leave it unstated. Three of the nine fire on an event instead of a date, and one source of the three, Future AGI, names a model version change as that event. Galileo names a red-flag signal beside its quarterly cycle, and AWS names a schedule of 60 to 90 days. Our own recommendation is the pairing this table keeps circling: a fresh sample, drawn on an event. See what judge calibration involves for the loop the cadence sits inside, and the validation protocol for where a recalibration run gets written down as a record rather than a one-off check.
What a band cannot tell you
Prevalence. A kappa is not comparable across samples with different label mixes. When one label takes almost every rating, chance agreement climbs and kappa falls for reasons that have nothing to do with the raters. Two panels can be equally good and land two bands apart. Our eval statistics hub works the arithmetic through; the short version is that the band is a reading of the sample as much as of the panel.
Task ambiguity. A rubric that genuinely admits two defensible answers puts a ceiling on agreement that no amount of rater training will lift. A low band on an ambiguous task is a finding about the rubric. That is the case where the number should send you back to the instructions rather than to a different judge, and it is how rubric drift starts to show up in the figures.
The cost of a wrong call. None of the five sources knows what your errors cost. A band that reads acceptable for low-stakes automation is written for low-stakes automation, and it stops being a reading the moment the judge gates something expensive. That is the judgment the tables cannot make for you, and it is the reason we publish them side by side instead of picking one. Reading a band against what a review actually costs is worked out on the review threshold optimizer.