LatentEval

INSTRUMENT | Judge reliability

Your judge disagrees with your humans. Which bias explains it?

Paste judge and human labels. See which bias mode explains the gap, ranked by how much each one accounts for, then follow the link to the calculator that tests it.

Paste a table of judge labels and human reference labels. This page ranks which bias mode accounts for most of the gap and links each row to the calculator that runs the confirmatory test. The bias checklist gives the run order and the pass line for every gate; this page reads your own labels and says which gate to run first. Where low agreement traces back to the rubric, the linter below names the structural flaw.

Judge and human labels

Paste a table with one row per eval item. The first two columns are the judge's label and the human reference. Additional columns can be assigned as metadata.

Check this value.

Rubric linter

Paste your rubric. One criterion per line, or a numbered or bulleted list.

Check this value.

Paste a rubric to see the findings.

G = J − H = (b − c) / nHow?

How this is calculated

The gap. Each item falls in one cell of a two-by-two table: a both pass, b judge pass and human fail, c judge fail and human pass, d both fail. Then J = (a + b) / n, H = (a + c) / n, and G = J − H = (b − c) / n. The two forms of G are the same number.

The ranking. For a mode whose stratifier splits the items into groups g: G_g = (b_g − c_g) / n_g for each group, and E = max(G_g) − min(G_g) for the mode. Modes are ranked by E descending, ties broken by position in the fixed list.

MIN_GROUP_N = 10
A group below this threshold is listed with its size but excluded from the spread calculation. A mode left with fewer than two usable groups is reported as not resolvable at this sample size.
MAX_GROUPS = 12
A column with more distinct values than this is refused as a stratifier and named with its count.
MAX_ROWS = 20,000
Above this the table is refused with the reason.
Bucketing rule for numeric columns
Rows carrying a value: 40 or more, quartiles; fewer than 40, a median split. Cut points are the type 7 quantiles (linear interpolation, the default in R and NumPy) of the values present. A value falling exactly on a cut point goes to the bucket at or below the cut, the same tie rule the median split uses. The rule in force is printed beside each mode.
Tie rule for the ranking
Two modes with equal E keep their declared order, so a rerun on the same paste ranks the same way.

A mode split into more groups tends to score higher because more groups mean more draws from the same underlying spread.

The linter. Ten checks, each resolving against a stated list or a stated pattern. Content words: case is normalized, the criterion is split on every character that is not a letter, the tokens in this stopword list of 66 words are dropped, and a trailing "ly" and then a trailing "s" are stripped wherever at least three letters remain. a all an and any are as at be been before both but by can do does each for from had has have if in into is it its least more most must no not of on or other should so some such than that the their then there these they this those to up was were when where which while will with would you your Two criteria overlap when the Dice coefficient over those sets, 2 × |A ∩ B| / (|A| + |B|), reaches 0.5.

The bias modes this page ranks come from our taxonomy of judge biases and the tests that detect them.

This page computes the gap and the ranking and stops there. Each row links to the calculator that owns the confirmatory test for that bias: the position bias calculator runs the swap analysis, the judge calibration calculator runs the calibration curve, the judge bias correction calculator runs the adjusted pass rate, the multi-rater agreement calculator runs the panel coefficients, and the inter-rater reliability calculator runs the baseline agreement. No coefficient and no p-value is computed on this page.

To track how these numbers move over time, log each run in the agreement tracker. To assemble the full validation record for a judge, use the validation report builder. For guidance on which published threshold bands to cite, see where threshold recommendations come from and when to report two.

The gap, the ranking and the constants above are the full specification. Every number on this page is reproducible from them by hand. The bucketing rule and the minimum group size are judgment calls: they decide when a group counts and how a numeric column is split, and both are stated here so you can see what they are rather than discovering them in a result that looks wrong.

Formula: G = J − H = (b − c) / n

Questions

Why does this page not just run the test itself?

Each bias has its own test with its own assumptions, sample-size requirements and interpretation. A single page that tried to run all five would either get the tests wrong or bury the assumptions. This page tells you which test to run first and hands the data to the calculator that does it right.

Why is a mode with more groups ranked higher?

The spread E is the distance between the highest and lowest group gap. More groups mean more draws from the same underlying distribution, so a wider spread is expected by chance alone. No correction is applied because the ranking is a triage order, not a hypothesis test. Compare two modes with the same group count, or follow the link to the confirmatory calculator where the group count enters the test.

My totals agree but the table still ranks a mode. Which is right?

Both. The overall gap is zero when opposite subgroup gaps cancel each other out, and the ranked table is showing you exactly where that cancellation happens. A zero total gap with a non-zero spread means the judge is biased in different directions in different subgroups, which is more informative than the total alone.

Why was my paste refused for one bad row?

One unreadable row means the page does not know what that item's labels are. Dropping it silently and computing over the rest would give you an answer built on an unknown subset of your own table, with no way to tell whether the dropped row changes the result. The refusal names the row so you can fix it and paste again.

Does anything I paste leave the browser?

No. The label table and the rubric are both read and processed entirely in this page. Nothing is sent to a server, stored, or shared.