Observed agreement
70.0%
INSTRUMENT | Judge reliability
4 cited sources
Cohen's kappa from the four counts of a 2x2 table: chance-corrected agreement between two raters, with a confidence interval and an interpretation band. Built for an LLM judge against human labels.
Cohen's kappa is chance-corrected agreement between two raters. Enter the four counts of a 2x2 table, say an LLM judge against human labels, for kappa with a confidence interval and its Landis and Koch band. A third rater sends you to the calculator for three or more raters. How near-perfect agreement can still collapse to low kappa is worked through in our statistics reference for eval numbers; the hub on judge reliability sets it beside calibration and bias.
Showing your last valid result. Update the inputs above to recompute.
Cohen's kappa (κ)
0.40
95% CI 0.15 to 0.65 | Fair agreement
Observed agreement
70.0%
Chance agreement
50.0%
Rates behind the score
Wilson score intervals at 95% coverage | 50 paired items
| Metric | Rate | Confidence interval |
|---|---|---|
| Raw agreement | (a + d) / n | 70.0% | 56.2% to 80.9% |
| Judge pass rate | (a + b) / n | 50.0% | 36.6% to 63.4% |
| Reference pass rate | (a + c) / n | 60.0% | 46.2% to 72.4% |
κ = (po - pe) / (1 - pe)How?
This tool reads one eval run as a 2x2 agreement table. Every item was scored pass or fail by the LLM judge and, independently, pass or fail by a human reference. Cell a counts items both passed, b the judge passed but the reference failed, c the reverse, and d items both failed. n is their sum.
Observed agreement is the share the two graders labeled the same way,
po = (a + d) / n. Two graders also agree on some items by chance
alone, at a rate set by how often each says pass. Chance agreement combines
the marginal rates: pe = (judge pass)(reference pass) + (judge
fail)(reference fail).
Cohen's kappa rescales observed agreement against that chance floor, so it
gives no free credit for the easy calls both graders get right:
κ = (po - pe) / (1 - pe). Kappa is 1 at perfect agreement, 0 at
chance, and negative when the graders agree less than their pass rates
predict. When pe = 1, every item sits in one category and kappa
is undefined.
The kappa confidence interval uses the large-sample standard error from
Fleiss, Cohen, and Everitt (1969), the form standard statistics packages
report. The interval is κ ± z·SE, clamped to the theoretical
range of -1 to 1. It is a normal approximation: at small n or extreme pass
rates it grows rough, so weigh its width more than its exact endpoints. The
raw agreement rate and the two marginal pass rates each carry a Wilson score
interval, which stays inside 0 to 1 and behaves on small counts.
Worked example. Take 50 items with a = 20, b = 5, c = 10, d = 15. Observed agreement is (20 + 15) / 50 = 0.70. The judge passed 25 of 50 (50%) and the reference passed 30 of 50 (60%), so chance agreement is 0.50·0.60 + 0.50·0.40 = 0.50. Kappa is (0.70 - 0.50) / (1 - 0.50) = 0.40, which the convention labels Fair. The standard error is 0.127, so the 95% interval is 0.40 ± 1.96·0.127, or 0.15 to 0.65. Raw agreement of 70% carries a 95% Wilson interval of 56% to 81%.
Agreement is one axis of judge reliability. For calibration and bias, see the LLM-as-a-judge reliability hub. Once you have a kappa, log it in the agreement tracker so this run has a series to sit inside.
Formula: κ = (po - pe) / (1 - pe)
No cutoff is universal. The Landis and Koch bands call 0.61 to 0.80 substantial and above 0.80 almost perfect, and many eval teams want substantial or better before a judge gates a release. Other published sources read the same coefficient differently; the full table, with the provenance behind each band lays out where they diverge. Always report the interval next to the point value so a reader sees the uncertainty.
It tells you how much of the agreement between two raters, whether that is two judges, two human graders, or a judge against a reference set, survives once the agreement their own pass rates would produce anyway is taken out. Kappa is 1 when the two label every item the same way, 0 when they agree exactly as often as chance predicts, and negative when they agree less often than that. It says nothing about who is correct: two raters can agree perfectly on the wrong label and still score a kappa of 1.
Count the four cells of the 2x2 table: both pass, both fail, and the two
ways the raters split. Observed agreement po is the share on
the diagonal. Chance agreement pe multiplies the two raters'
pass rates and adds the product of their fail rates. Kappa is
(po - pe) / (1 - pe). The "How this is calculated" panel
above runs that arithmetic on a 50-item example.
Percent agreement is the plain share of items the two graders labeled the same way. It counts the easy items both graders get right, which inflates the number on skewed pass rates. If a judge and a reference each pass 95% of items, they land on the same label about 90% of the time by chance alone. Kappa subtracts that floor, so the score reflects agreement beyond chance.
You do not need a spreadsheet. Excel ships no kappa function, so the usual route is a hand-built grid of COUNTIF formulas for the four cells and then the marginal products by hand, where one transcription slip changes the answer silently. Enter the same four counts here and you also get the standard error and the confidence interval, which the spreadsheet recipe leaves out.
Negative kappa means the judge and reference disagree more than their individual pass rates predict. In practice it usually points to an inverted label, a prompt that flips the judge's polarity, or a reference set scored under different instructions. Inspect the off-diagonal cells b and c.
More items narrow the interval. The standard error here is a large-sample approximation, so below roughly 30 paired items the kappa interval is a rough guide. When you need a target precision, size the sample from the interval width you can tolerate on the underlying pass rate.
This is Cohen's kappa for two raters on a binary label. Fleiss' kappa generalizes it to three or more raters, and the Fleiss' kappa calculator for three or more raters computes it alongside Krippendorff's alpha and the intraclass correlation coefficient (ICC). For an ordinal or graded label, weighted kappa credits near-misses; this tool uses the unweighted binary form and does not compute the linear or quadratic weighted variants.