LatentEval

INSTRUMENT | Judge reliability

Fleiss' Kappa, Krippendorff's Alpha and ICC Calculator

12 cited sources

Agreement among three or more raters or LLM judges: Fleiss' kappa, Krippendorff's alpha, Gwet's AC1 and intraclass correlation, with percent agreement, bootstrap intervals, and per-rater diagnostics.

Fleiss' kappa measures agreement among three or more raters; with exactly two it collapses to Scott's pi, so use the two-rater Cohen's kappa calculator. Paste an items-by-raters matrix for Fleiss' kappa, Krippendorff's alpha, ICC, Gwet's AC1, and per-rater diagnostics. AC1 is the coefficient that holds up when one label dominates. Agreement sets a floor on trust without proving correctness; what counts as good enough is a convention, and the eval statistics hub maps which number needs which test.

One item per line, one rater per column, separated by commas, tabs, or spaces. Leave a cell blank or write NA for a missing rating (with spaces as the separator, write NA). A first line of distinct, non-numeric rater names is read as a header when none of them appears in the data below. Labels are compared exactly. Up to 25 raters and 5000 items.

Check this value.

Scale and interval

Nominal for labels with no order such as pass, fail, partial. Ordinal for ranked scores where a near miss is a smaller disagreement, such as a 1 to 5 rubric. Interval when the gaps between scores are meaningful. Ordinal and interval need numeric labels.

Check this value.

For the bootstrap intervals on kappa and alpha and the F-based ICC intervals.

Check this value.

The rubric points a rater could have chosen, comma separated. This changes Gwet's AC1 only: a point nobody used still counts toward its chance term, which no other coefficient on this page reads. Leave it empty to use the labels present in the data.

Check this value.

Krippendorff's alpha (ordinal)

0.851

95% bootstrap interval 0.706 to 0.896 | 20 items | 4 raters | 5 categories | 3 missing cells

On the Landis and Koch scale, Krippendorff's alpha of 0.851 reads almost perfect and Gwet's AC1 of 0.430 reads moderate. Alpha clears the 0.80 bar Krippendorff sets for reliable conclusions. That scale comes from Landis and Koch 1977, via AHRQ, read 2026-08-21, and was written for observer agreement on categorical medical data. The 1977 paper was written for kappa. The bands are a convention, and the published mappings disagree with each other.

Fleiss' kappa

0.365

Observed agreement

53.2%

Agreement coefficients on the 20 items, 4 raters

CoefficientEstimateIntervalBasis
Krippendorff's alpha, nominal 0.413 20 items, 77 ratings
Krippendorff's alpha, ordinal (selected) 0.851 0.706 to 0.896 20 items, 77 ratings
Krippendorff's alpha, interval 0.840 20 items, 77 ratings
Fleiss' kappa 0.365 0.196 to 0.496 17 complete items
Fleiss' kappa z against chance 7.19 null SE 0.051
Gwet's AC1, unweighted 0.430 0.291 to 0.586 20 items with 2+ ratings, 5 categories observed
Observed agreement, exact match 53.2% per-item pair agreement, weighted by ratings
Chance agreement 20.3% two ratings drawn from the pooled label mix, without replacement, over every pairable item
Chance agreement, Gwet basis 19.6% prevalence term over 5 categories; Fleiss' term on the complete items is 21.2%

Intraclass correlation on the 17 complete items with 95% F-based intervals; ICC(1) one-way random, ICC(2) two-way random with absolute agreement, ICC(3) two-way mixed with consistency

FormICCIntervalF (df)p
ICC(1), single rater 0.839 0.704 to 0.930 21.80 (16, 51) <0.001
ICC(1,k), mean of raters 0.954 0.905 to 0.981 21.80 (16, 51) <0.001
ICC(2,1), single rater 0.841 0.657 to 0.936 33.24 (16, 48) <0.001
ICC(2,k), mean of raters 0.955 0.885 to 0.983 33.24 (16, 48) <0.001
ICC(3,1), single rater 0.890 0.788 to 0.953 33.24 (16, 48) <0.001
ICC(3,k), mean of raters 0.970 0.937 to 0.988 33.24 (16, 48) <0.001

Label distribution over all 77 ratings, with Fleiss' per-category kappa on the 17 complete items

CategoryRatingsShareCategory kappa
1 8 10.4% 0.339
2 16 20.8% 0.207
3 19 24.7% 0.346
4 19 24.7% 0.345
5 15 19.5% 0.588

Per-rater diagnostics on the 20 items: items rated, how often the rater matches the other raters' majority label, Krippendorff's ordinal alpha with that rater's column removed, and the rater's most-used label

RaterRatedMatches majorityAlpha withoutChangeMean biasTop label
judge_a 19 68.8% of 16 0.804 -0.047 -0.17 2 (21%)
judge_b 20 50.0% of 20 0.881 +0.030 -0.30 4 (30%)
judge_c 19 38.9% of 18 0.898 +0.047 +0.61 5 (32%)
human 19 64.7% of 17 0.824 -0.027 -0.12 2 (26%)

Pairwise linear-weighted Cohen's kappa for each rater pair on the items both rated, with exact-match agreement in parentheses

Raterjudge_ajudge_bjudge_chuman
judge_a 0.72 (58%) 0.59 (39%) 0.89 (83%)
judge_b 0.72 (58%) 0.51 (42%) 0.62 (47%)
judge_c 0.59 (39%) 0.51 (42%) 0.59 (44%)
human 0.89 (83%) 0.62 (47%) 0.59 (44%)
Export

The first line was read as rater names (judge_a, judge_b, judge_c, human) and is not counted as an item. If that line was meant as data, add a line of rater names above it.

3 of 80 cells are missing. Krippendorff's alpha, Gwet's AC1, the pairwise matrix and the rater diagnostics use every rating present (20 of 20 items carry at least two ratings). Fleiss' kappa and the ICC use the 17 items every rater scored, so 3 items are left out of those figures.

Fleiss' kappa treats the scale as unordered categories, so a rating one step away counts as full disagreement. That is why it sits below the ordinal alpha, which credits near misses.

Intervals on kappa and alpha are 95% percentile bootstrap intervals from 1000 item resamples (seed 20260816, so a rerun reproduces them).

The ICC reads the labels as numeric scores on the 17 complete items. ICC(2,1) is the figure for a single rater drawn from a larger pool when absolute agreement matters; ICC(3,1) when these exact raters are the panel and only consistency matters; the k forms rate the panel average.

α = 1 − D_o / D_e; κ = (P̄ − P̄_e) / (1 − P̄_e)How?

How this is calculated

The input is a matrix with one row per judged item and one column per rater. Raters can be LLM judges, human annotators, or a reference set treated as one more rater, so the same figures serve inter-rater and inter-coder reliability, whether the panel is a jury of LLM judges from different model families or a set of human coders working one rubric. Cells may be empty. Each coefficient is computed on the slice of the data its definition allows, and the coefficient table names that slice next to every figure, because a coefficient quoted without its basis cannot be compared with anyone else's.

Krippendorff's alpha, the headline. Every item carrying at least two ratings contributes its within-item pairs of values to a coincidence matrix; alpha is one minus the observed disagreement in that matrix divided by the disagreement expected if the same pool of values were shuffled across items (Krippendorff 2011). Missing cells simply contribute fewer pairs, which is why alpha is the one coefficient here that needs no rows dropped. The level of measurement sets the difference function: nominal counts any mismatch as one, ordinal uses Krippendorff's rank distance measured through the coincidence marginals, so a 4 against a 5 disagrees less than a 1 against a 5, and interval uses the squared difference of the numeric values. Ordinal and interval need numeric labels; unused points on the scale do not change the ordinal figure. Krippendorff's own reading is that 0.80 supports reliable conclusions and 0.667 supports tentative ones; the Landis and Koch bands (poor, slight, fair, moderate, substantial and almost perfect) are shown as well, being the more familiar convention, with the caveat that both are conventions rather than tests.

Fleiss' kappa. Fleiss (1971) generalizes Scott's pi to any number of raters and needs the same number of ratings on every item, so it runs on the complete-case items, those every rater scored. Its observed agreement is the mean share of agreeing rater pairs per item; its chance agreement is the sum of squared pooled label shares; kappa is their difference scaled by the agreement chance leaves available. The observed and chance rows in the coefficient table are the versions defined on every pairable item, weighting each item by its number of ratings and drawing the chance pair from the pooled labels without replacement; on complete data they equal Fleiss' figures, and on any data one minus the ratio of their complements is the nominal alpha, so those two rows explain the nominal alpha rather than the kappa. Fleiss' kappa treats every scale as unordered, which is why it lands well below the ordinal alpha on rubric scores. The per-category kappa in the label table is Fleiss' own decomposition and shows which labels the panel agrees on. The z row tests kappa against zero using the Fleiss, Nee and Landis (1979) variance under the null; it says whether agreement is detectably better than chance, which is a weak bar, so the bootstrap interval is the figure to quote.

Two raters, and the sibling tool. With exactly two raters, Fleiss' kappa is Scott's pi: it pools both raters' label rates into one marginal. Cohen's kappa keeps each rater's own rates and is the figure the two-rater kappa calculator reports for a judge against a human on pass/fail labels, with its large-sample interval. The two agree only when both raters use the labels at the same rates. The pairwise matrix on this page is Cohen's kappa for every rater pair, so a two-column nominal paste reproduces the sibling's number in the matrix cell while the headline stays the multi-rater family. Two raters on a binary label belong in the sibling tool; three or more raters, more than two labels, ordered scores, or gaps in the matrix belong here.

Intraclass correlation. When the labels are numeric and the level is ordinal or interval, a two-way ANOVA on the complete items gives the six Shrout and Fleiss (1979) forms, in their notation; McGraw and Wong (1996) write ICC(2,1) as ICC(A,1) for absolute agreement and ICC(3,1) as ICC(C,1) for consistency. ICC(1) treats each item's raters as a random draw with no rater identity. ICC(2,1) treats raters as a random sample from a larger pool and asks for absolute agreement, so a rater who scores everything one point higher is penalized; that fits a judge prompt you might swap for another. ICC(3,1) treats these exact raters as the panel and asks only for consistency, so a constant offset is forgiven; that fits a fixed ensemble whose scores you will average and rank. The k forms rate the reliability of the panel mean rather than one rater. Intervals are the F-based ones from those papers; the absolute-agreement interval uses the Satterthwaite degrees of freedom McGraw and Wong give, while the F test in the table uses the ordinary ANOVA degrees of freedom shown beside it. Fleiss' kappa is still computed on numeric data, and the gap between it and the ICC is the price of ignoring the order of the scores.

Rater diagnostics and the pairwise matrix. For each rater: how often its label matches the plurality of the other raters (items where the others tie are skipped); alpha recomputed with that rater removed and the change against the full panel, so a positive change names a rater the panel agrees better without; the mean signed gap between the rater's score and the mean of the others' scores on the same items, which reveals a lenient or harsh judge on numeric scales; and the rater's most-used label with its share, which reveals a rater who says pass to nearly everything. The pairwise matrix is Cohen's kappa for each pair on the items both rated, unweighted at the nominal level, linear-weighted on rank distance at the ordinal level, and quadratic-weighted on value distance at the interval level, with the exact-match share in parentheses. A rater whose row is uniformly low disagrees with everyone; a single low cell is a pair-specific problem.

Why kappa can look bad when agreement looks good. Chance-corrected coefficients divide by the agreement chance leaves available. When one label dominates, say 92% of all ratings are pass, chance agreement is already high and even 95% raw agreement leaves a modest kappa; when two raters disagree about how often to use a label, the coefficient moves again. Feinstein and Cicchetti (1990) named these the prevalence and bias paradoxes. Neither is a defect in the coefficient: it is reporting that most of the raw agreement was cheap. The label table shows the prevalence, the coefficient table shows raw and chance agreement side by side, and the honest report quotes all three.

Gwet's AC1. AC1 answers the same question with a different chance term. Fleiss asks how often two raters would agree by picking labels at the rates the panel used overall, which climbs toward one as a single label takes over. Gwet asks how often two raters would agree by chance on the items that are genuinely hard to call, so a lopsided label mix pushes his chance term down instead of up. In numbers, the term is the sum of p(1 − p) across the scale points, divided by one less than the number of points. AC1 runs on every item with at least two ratings, the same set Krippendorff's alpha uses, so a missing cell costs it nothing. Read on the same items, the two chance terms meet and the two coefficients are equal when every label is used equally often, and away from that point AC1 is the higher of the two. Fleiss' kappa drops the partly rated items that AC1 keeps, so on a matrix with missing cells the two are reading different item sets and AC1 can come out below. The unweighted form is what this page reports at every level. One input moves AC1 and nothing else: declaring the rubric's scale points. A five-point rubric where raters only ever used 4 and 5 has a chance term computed over two points unless you say there were five, and the coefficient table names which count it used.

Uncertainty. Kappa and alpha carry percentile bootstrap intervals from item resamples. For nominal coefficients, Zapf and colleagues (2016) found that interval to hold its coverage where the closed-form variances did not; the same item-level resampling is applied at the ordinal and interval levels, where it is the standard practical choice rather than a separately validated one. The resampling is seeded, so a rerun on the same data reproduces the same interval, and resamples on which the coefficient is undefined are dropped and counted. The interval widens as items get scarce; with fewer than about ten items it is only indicative, and with a single item none is shown. The ICC intervals are the analytic F-based ones.

Worked example. The preloaded panel is three LLM judges and one human scoring 20 items on a 1 to 5 rubric with three cells missing. At the ordinal level alpha is 0.851 (95% bootstrap interval 0.706 to 0.896), which clears Krippendorff's 0.80 bar; the interval alpha is 0.840 and the nominal alpha 0.413, the same data read as unordered labels. Fleiss' kappa on the 17 complete items is 0.365, and observed exact-match agreement is 53.2% against 20.3% by chance. ICC(2,1) is 0.841 and ICC(3,1) is 0.890; the gap is the cost of absolute agreement, and its source shows in the rater table: judge_c scores 0.61 points above the rest of the panel on average and matches the others' majority on only 39% of items, and removing judge_c would raise alpha to 0.898. Switch the level to nominal and the same panel reads as fair agreement, because every one-point difference now counts as a full miss.

Honest limits. Every figure is arithmetic on the matrix you paste. Agreement among judges is a floor on trust, not a proof of correctness: a panel of judges can agree with one another and still be wrong in the same direction, which is the failure mode worked through in testing whether an LLM judge is reliable and in the catalog of LLM judge biases. Kappa and alpha also depend on the label mix in the sample, so compare panels only on the same items and quote the coefficient with its level, its basis, and its interval, the way the judge validation report expects it.

Formula: α = 1 − D_o / D_e; κ = (P̄ − P̄_e) / (1 − P̄_e)

Questions

I have two raters on pass/fail labels. Which tool?

The Cohen's kappa calculator for two raters. It takes the 2x2 table directly and reports Cohen's kappa with its large-sample interval, plus the marginal pass rates. This page accepts two raters as well, and its pairwise cell reproduces that Cohen's kappa, but its headline is Krippendorff's alpha and its Fleiss' kappa is Scott's pi, which pools the two raters' label rates. Use this page once you have a third rater, a third label, an ordered scale, or gaps in the matrix.

What is a good Fleiss' kappa or Krippendorff's alpha value?

There is no test to pass, only conventions, so quote the number with its interval and its basis. Two conventions are in common use and both are shown with the result here. Krippendorff's own reading is that 0.80 supports reliable conclusions and 0.667 supports tentative ones. The other is the Landis and Koch scale, whose bands read poor, slight, fair, moderate, substantial and almost perfect, the last of them printed as 0.81 to 1.0; those edges and their source are published in full in the table of kappa thresholds and where each one comes from. Neither convention survives a change in label mix. The same panel scoring a rarer label will land lower for reasons that have nothing to do with the raters, which is why the bootstrap interval and the item count belong next to the coefficient wherever it is quoted.

How do the published kappa bands compare?

Five mappings from four sources, side by side. They disagree: 0.65 is substantial on one scale and moderate on another, and one of them attaches a deploy decision to each band instead of a word. Each row names who published it and the day we read it. None of them is a standard, and a band written for one coefficient does not transfer to another, which is why this calculator only reads a set that names the coefficient it computed. The full table, with the calibration cadences that go with it, is published separately.

Published interpretation bands for chance-corrected agreement coefficients

RangeWhat the source says
Landis and Koch 1977, via AHRQ, read 2026-08-21. Written for Cohen's kappa, Fleiss' kappa, Gwet's AC1, Krippendorff's alpha, either pairing.
< 0 Poor
0.0 to 0.20 Slight
0.21 to 0.40 Fair
0.41 to 0.60 Moderate
0.61 to 0.80 Substantial
0.81 to 1.0 Almost perfect
McHugh 2012, Biochemia Medica, read 2026-08-21. Written for Cohen's kappa, human against human.
0 to .20 None, 0 to 4 percent of data reliable
.21 to .39 Minimal, 4 to 15 percent
.40 to .59 Weak, 15 to 35 percent
.60 to .79 Moderate, 35 to 63 percent
.80 to .90 Strong, 64 to 81 percent
above .90 Almost perfect, 82 to 100 percent
AWS sample-GEDD, read 2026-08-21. Written for Cohen's kappa, judge against human.
< 0.00 Rubric is inverted, fix immediately
0.00 to 0.20 Judge is unreliable, do not use
0.21 to 0.40 Major rubric revision needed
0.41 to 0.60 Usable with human review on flagged cases
0.61 to 0.79 Acceptable for low-stakes automation
>= 0.80 Deploy autonomously in CI
Future AGI, read 2026-08-21. Written for Cohen's kappa, Fleiss' kappa, human against human.
below 0.4 The rubric is ambiguous, rewrite it
0.4 to 0.6 Weak, the rubric is tunable
above 0.6 Acceptable
above 0.8 Strong rubric
Future AGI, read 2026-08-21. Written for Cohen's kappa, Fleiss' kappa, judge against human.
above 0.6 Acceptable for production
above 0.8 Strong
How many raters do I need for Fleiss' kappa?

Three is the minimum this page is built for, and Fleiss' kappa is defined for three or more. With exactly two raters kappa reduces to Scott's pi, which pools both raters' label rates, so a two-rater job belongs in the sibling tool. Adding raters buys a steadier estimate; the width of the interval is driven by the item count, because the bootstrap resamples items and not raters. Below about ten items the interval is only indicative. The matrix accepts up to 25 raters and 5000 items.

Fleiss' kappa or Krippendorff's alpha: which should I use?

Krippendorff's alpha, with the level named, is the more general figure: it handles any number of raters, missing cells, and ordered scales under one definition, and on complete nominal data it lands within a small-sample correction of Fleiss' kappa. Fleiss' kappa is the number many readers expect, so quoting both costs nothing. If the panel scores are numeric and you will average them, the ICC is the figure that matches how the scores will be used. Whatever you quote, attach the basis: how many items and ratings went in, and the interval.

Why is Fleiss' kappa so much lower than alpha on my 1 to 5 scores?

Fleiss' kappa treats the five scores as unordered labels, so a 4 against a 5 is as much a disagreement as a 1 against a 5. Ordinal and interval alpha credit the near miss. On a rubric where raters mostly land within a point of each other, that difference is large, and the ordinal alpha is the fairer read of the panel. Switch the level to nominal and alpha drops to meet kappa, which is a useful check that the gap is the scale and not the data.

Percent agreement is 92% but kappa is 0.35. Which is right?

Both. Look at the label table: when one label holds most of the ratings, chance agreement is already high, and kappa reports how much of the raw agreement was beyond that. A panel that says pass to almost everything agrees with itself for cheap. There is a third figure for exactly this case. Gwet's AC1 uses a chance term that falls as one label takes over, so on a 95% pass set it can read close to one where kappa reads moderate, and the gap between them measures the label mix. Report raw agreement, chance agreement, and the coefficient together; if the rare label is the one you care about, read the per-category kappa for it, and see what the published bands say a value of that size is worth before treating either number as a pass mark.

Should I report Gwet's AC1 or Fleiss' kappa?

One figure, with its interval and the items it ran on. A report carrying five coefficients invites the reader to pick the flattering one, which is the case Agreement Metrics for LLM-as-Judge Evaluation makes; that paper picks Cohen's kappa. Pick before you look at the numbers, and pick on the data. If one label holds most of your ratings, kappa is reporting a real property of that sample and AC1 is the steadier estimate of how the raters behave, so quote AC1 and show the kappa beside it so nobody thinks it was hidden. If the labels are reasonably balanced, the two nearly agree and kappa is the figure more readers already know. Either way, name the coefficient, the level, the item count, and the interval.

Which intraclass correlation (ICC) form should I report for a panel of LLM judges?

If each judge is one prompt or model you might swap for another and a systematically lenient judge should count against reliability, ICC(2,1) is the figure for a single judge's score and ICC(2,k) for the panel mean. If these exact judges are the fixed ensemble and you only need their scores to move together, ICC(3,1) and ICC(3,k). The two families differ by whether a constant offset between judges is a disagreement, and the mean-bias column in the rater table shows how large that offset is.

How are missing ratings handled?

Krippendorff's alpha, the pairwise matrix, and the rater diagnostics use every rating that is present; an item with a single rating contributes nothing to alpha because it has no pair. Fleiss' kappa and the ICC need the same raters on every item, so they use only the items every rater scored, and the note under the results says how many were left out. If most cells are missing, alpha is the only figure to trust, and its interval will say how little the data support.

Does high agreement mean the judges are accurate?

No. Agreement measures whether the raters see the same thing, not whether what they see is right. Judges built on similar models share blind spots and can agree confidently on a wrong verdict. Treat agreement as a gate: a panel that cannot agree with itself cannot be trusted, but one that can still needs checking against ground truth, which is what the judge bias correction calculator and the position bias calculator are for.

Sources

  1. Fleiss (1971), Measuring Nominal Scale Agreement Among Many RatersPsychological Bulletin Retrieved
  2. Fleiss, Nee and Landis (1979), Large Sample Variance of Kappa in the Case of Different Sets of RatersPsychological Bulletin Retrieved
  3. Krippendorff (2011), Computing Krippendorff's Alpha-ReliabilityUniversity of Pennsylvania, Annenberg School for Communication Retrieved
  4. Shrout and Fleiss (1979), Intraclass Correlations: Uses in Assessing Rater ReliabilityPsychological Bulletin Retrieved
  5. McGraw and Wong (1996), Forming Inferences About Some Intraclass Correlation CoefficientsPsychological Methods Retrieved
  6. Cohen (1968), Weighted Kappa: Nominal Scale Agreement with Provision for Scaled Disagreement or Partial CreditPsychological Bulletin Retrieved
  7. Feinstein and Cicchetti (1990), High Agreement but Low Kappa: The Problems of Two ParadoxesJournal of Clinical Epidemiology Retrieved
  8. Zapf, Castell, Morawietz and Karch (2016), Measuring Inter-Rater Reliability for Nominal Data: Which Coefficients and Confidence Intervals Are Appropriate?BMC Medical Research Methodology Retrieved
  9. Gwet (2008), Computing Inter-Rater Reliability and Its Variance in the Presence of High AgreementBritish Journal of Mathematical and Statistical Psychology Retrieved
  10. Wongpakaran, Wongpakaran, Wedding and Gwet (2013), A Comparison of Cohen's Kappa and Gwet's AC1 When Calculating Inter-Rater Reliability CoefficientsBMC Medical Research Methodology Retrieved
  11. Hartling and colleagues (2012), Table 2: the Landis and Koch 1977 reading of kappa, as the AHRQ methods guide prints itAgency for Healthcare Research and Quality, via NCBI Bookshelf Retrieved
  12. McHugh (2012), Interrater Reliability: The Kappa StatisticBiochemia Medica Retrieved