LatentEval

Glossary

Krippendorff's alpha (eval agreement)

Krippendorff's alpha is the proportion by which a set of labels falls below the disagreement chance would have produced, computed for any number of raters, on any measurement scale, with missing labels tolerated. Eval teams use it for judge-against-human agreement on graded rubrics.

Krippendorff’s alpha is a chance-corrected agreement coefficient, defined as one minus the ratio of observed disagreement to the disagreement expected if the same pool of labels had been scattered across items at random. It computes over any number of raters, on any measurement scale, with missing labels left in place instead of dropped. Applied to an eval, it scores how far a judge’s labels reproduce a human reference beyond what coincidence would supply, on designs that outgrow Cohen’s kappa and its limit of exactly two raters labeling every item on one nominal scale.

Because alpha divides one estimated disagreement sum by another, it carries no dependable closed-form standard error, so its sampling distribution has to be generated from the reliability data itself. Hayes and Krippendorff bootstrap the pairable judgments and read the interval off the 2.5th and 97.5th percentiles: in their worked example of five observers scoring 40 units, ordinal alpha came to 0.7598 with a 95% interval of 0.7078 to 0.8078 across 10,000 resamples (Communication Methods and Measures, 2007, peer-reviewed).

Agreement is one of three axes a judge clears before its scores license a decision, sitting alongside calibration and bias in the judge-reliability frame. Alpha carries that axis once a design outgrows two raters and a binary label, and its interval comes from the same bootstrap resampling machinery behind every other interval an eval score has to carry here. A lone alpha with nothing bracketing it describes one labeling session and licenses nothing past it.

How to calculate Krippendorff’s alpha

Lay the reliability data out as a matrix with one row per item and one column per rater, leaving a cell blank wherever a rater never saw that item. Collapse it into a coincidence matrix, which counts every pairable value inside a unit in both directions and divides each unit’s pairs by the number of raters who scored it, less one. Choose the difference function that matches your scale: nominal treats every disagreement as equally bad, ordinal charges by rank distance, interval by squared difference. Observed disagreement is the weighted total those differences produce, expected disagreement is what the same values would produce shuffled across units at random, and alpha is one minus their ratio, so 1 is perfect reproduction and 0 says the labels carry no information about the items (Krippendorff, Computing Krippendorff’s Alpha-Reliability).

Report a bootstrap interval alongside the point estimate, declare the minimum alpha you will accept before the labels come in, and afterwards report the probability that the true value falls below that minimum. In the worked example above, that probability was 0.0125 against a floor of 0.70 and 0.9473 against a floor of 0.80, so one coefficient produced two opposite verdicts on a threshold chosen in advance.

Compute it on the human-labeled calibration slice rather than the full suite the judge grades alone. Recompute it once per rubric dimension, since alpha evaluates reliability one variable at a time and a figure blended across helpfulness, factuality and formatting hides which dimension the judge is fumbling. Our inter-rater reliability calculator covers the two-rater binary case with Cohen’s kappa and a confidence interval; alpha’s general form wants a statistics package behind it.

Krippendorff’s alpha vs Cohen’s kappa

Cohen’s kappa corrects raw percent agreement between exactly two raters on nominal categories, dividing the excess of observed agreement over chance by the room left above chance. Its chance term is what those two raters would hit if their labeling habits were statistically independent of one another.

That baseline is the whole source of the divergence.

Alpha builds its expected disagreement from the pooled distribution of all values across all raters, which keeps the coefficient tied to the data whose reliability is in question. On identical labels the two coefficients can move in opposite directions. When a judge and a human grader converge on the same overall pass rate, kappa reads their shared marginal distribution as inflated chance agreement and marks them down for it, which Hayes and Krippendorff describe as “punishing observers for agreeing on the frequency distribution of categories”. A systematically stricter grader does the reverse, widening the gap between the two marginals, lowering kappa’s chance term and lifting the coefficient on labels that are genuinely less reliable.

A structural limit sits underneath the statistical one. Cohen’s kappa in its original two-rater, unweighted form has no definition for three raters or for an ordered rubric, and the family answers each of those gaps on its own terms: weighted kappa charges by rank distance across ordered categories, and Fleiss’s generalization extends the chance term past two raters while staying on nominal categories. What alpha adds is coverage in a single coefficient. A panel of five graders working a 1-to-5 scale, with blanks wherever a grader never saw an item, is one alpha computation on the ordinal metric, and no combination of the kappa variants takes arbitrary rater counts, an ordered scale and missing judgments together.

Quote kappa where your setup is literally two raters on a binary label and your reviewers expect the coefficient they know. Alpha is the one to reach for where the design has more raters, a graded scale, or gaps in coverage. Publishing both, where both are computable, makes the marginal effect visible.

The scale you declare changes the number

Alpha’s difference function is a modeling decision, and it moves the answer by more than most reporting practice admits. Hayes and Krippendorff scored one dataset four ways and got 0.4765 treating the judgments as nominal, 0.7598 as ordinal, 0.7574 as interval, and 0.6621 as ratio. The labels and the raters were identical across all four, and the coefficients span nearly 0.3.

An LLM judge scoring a 1-to-5 rubric is exactly this case. Collapse those grades to unordered categories and a 4 against a 5 counts as the same failure as a 1 against a 5, which drags the coefficient toward that nominal figure. Declare the ordinal metric and the near miss costs a fraction of what the far one does. Both readings are defensible, so state which metric produced the number: a bare alpha in a model card cannot be read without it.

Content analysis and eval agreement ask alpha different questions

Alpha comes from content analysis, where it gates a study: a team writes a codebook, several human coders apply it to the same sample, and analysis proceeds once the coefficient clears an agreed threshold. That meaning is legitimate, and most of the literature carrying this name is about it. An eval borrows the arithmetic and changes two things underneath.

One side of the comparison is a reference rather than a peer. Alpha treats raters as freely interchangeable, so a low value says the labeling process is unreliable without saying whose labels are wrong, and an eval does want to know: the human slice is the standard the judge is held to. Reading a low alpha as a verdict on the judge takes those human labels on trust.

Retraining is the other thing that does not carry across. A codebook producing poor agreement gets clarified, the coders re-run the sample, and the number moves. A judge model answers a rewritten rubric less predictably, and the same prompt can shift its effective standard across a long run, which is rubric drift and sits outside what an agreement coefficient can name.

Teams arriving from either field usually want one number: the alpha above which a judge is good enough. No field-wide constant is available to hand over, because the cutoff encodes what a mislabeled item costs downstream, and that cost differs between a nightly regression gate and a published leaderboard claim. Use the construction the coefficient’s own authors relied on. Pick your minimum before the labels exist, bootstrap the distribution, and report the probability of falling short of it; their example ran that against six candidate minima from 0.90 down to 0.50, publishing a threshold decision instead of inheriting one.

A judge that agrees with people and is still wrong has a calibration or a bias problem underneath, each with its own test. Alpha is one instrument in a larger kit, and the statistics behind a trustworthy eval maps the rest of it. Once the judge has been characterized, correcting a judge’s scores before you report them picks the number up and carries it as far as publication.