Glossary
Self-preference bias (LLM judges)
Self-preference bias is an LLM judge scoring its own model family's responses above the score a reference standard gives the same text, measured as the gap between the judge's own-family win rate and the rate human annotators or a cross-family panel assign to identical pairs.
Self-preference bias is the gap between the score an LLM judge gives responses from its own model family and the score a reference standard gives the same responses. Measuring it needs that reference, either human annotators who ranked the pairs or a panel of judges from other families reading identical text. It sits beside verbosity bias and position bias among the distortions that move a verdict with no change in answer quality, and its size shifts depending on which family you asked to grade. What each judge bias is called, and which test separates it from the next one, is set out term by term.
A judge scored in isolation produces no number at all, because the quantity is a gap between two readings of one artifact.
Because the quantity is a difference between two win rates computed on the same pairs, it inherits the uncertainty of both, and a gap quoted as a lone percentage is a point estimate of a difference with its uncertainty left off. Report it with an interval on the difference and the pair count that produced it, on the same reasoning that governs any other eval confidence interval. The pairs are matched, since the same items were scored twice, so the discordant pairs carry the information and a paired test is the right instrument to run on them.
The effect is documented and its size is not settled. Panickssery, Bowman and Feng define self-preference as an evaluator scoring “its own outputs higher than others’ while human annotators consider them of equal quality.” The same work reports “a linear correlation between self-recognition capability and the strength of self-preference bias,” with models including GPT-4 and Llama 2 showing non-trivial accuracy at telling their own text apart from another model’s (peer-reviewed, NeurIPS 2024). Recognition sits under the mechanism, which is why blinding provenance comes before anything else, and why a cross-family reader settles what blinding on its own cannot. Self-preference belongs to the bias axis of judge characterization, alongside agreement and calibration in the three-axis frame. Treat any ranking a model produced over its own family as unproven until a reader outside that family has scored the same pairs.
How to calculate self-preference bias
Assemble pairwise comparisons in which one side came from the judge’s own model family and the other did not, and strip provenance markers from both sides. Have the judge rank every pair in both presentation orders, so the slot each answer sat in does not ride along inside the number. Score those same pairs again with a reference: human annotators where the budget reaches, otherwise a panel of judges drawn from other families, checked for agreement past chance with the inter-rater reliability calculator. Self-preference is then the judge’s own-family win rate minus the reference’s own-family win rate on those pairs, in percentage points.
Both win rates go beside the gap, because a gap of 6pp sits differently at 52% against 46% than it does at 94% against 88%. Put a Wilson interval on each rate and a paired interval on the difference between them. Instrument this wherever your harness writes a judge verdict down, with the generating model recorded on every response, because an own-family win rate cannot be reconstructed once the provenance column has been dropped, and nothing else in a verdict row carries it.
Part of the gap can be legitimately earned. Chen and colleagues scored self-preference on verifiable benchmarks in mathematics, factual knowledge and code, where an answer is right or wrong independently of any judge. Much of a stronger model’s preference for its own output tracked answers that were in fact correct, and the harmful share concentrated on the items the model had got wrong.1 A nonzero gap therefore licenses an investigation rather than an automatic adjustment, and where an adjustment is warranted it is the bias-corrected reporting procedure run on a judge whose error rates you actually measured.
Self-preference bias vs verbosity bias
Verbosity bias is a judge preferring the longer of two responses that are equally correct, so its score tracks length. Self-preference runs on a different axis, favoring the judge’s own family’s text at whatever length that text happens to be. Model families have a house length and a house way of opening an answer, so the two effects arrive already tangled: a judge that rewards length will post a high own-family win rate whenever its siblings write long, and none of that gap came from recognition.
The confound is symmetrical, and each way round it spoils a different number. An uncontrolled self-preference measurement can report a gap that’s entirely length, which shows up when you hold the judge fixed, swap in equally long responses from another family, and watch the gap move. A verbosity measurement can be contaminated the other way, when the long side of your length-varied pairs happens to be the judge’s own family, so a family effect gets absorbed into the length coefficient and reported as length.
Separating them takes an explicit control instead of a caveat. The cheap version is stratification: bin the pairs by length difference and read the own-family gap inside each bin, so a gap that survives where the two sides are the same length isn’t doing length. The stronger version fits a model predicting the judge’s preference from length difference and family membership together, then reads each with the other held fixed. That is the move behind length-controlled AlpacaEval, where a generalized linear model fit on the biased annotator’s preferences was queried at a zero length difference, lifting Spearman correlation with Chatbot Arena from 0.94 to 0.98.2
A third published result complicates the pair further. Wataoka, Takahashi and Ri report that LLM judges “assign significantly higher evaluations to outputs with lower perplexity than human evaluators, regardless of whether the outputs were self-generated.”3 That places familiarity of the text to the judge underneath both effects rather than beside them. Report the two effects separately, and store the length distribution and the generating family for every response, so that a later analysis can still ask which of the three is doing the work.
Self-preference bias vs self-enhancement bias
Self-enhancement bias is the older name, borrowed from social psychology, where it describes a person’s motivated tendency to see themselves more favorably than the evidence supports. The MT-Bench authors applied that name to models. Zheng et al. observed GPT-4 rating its own answers with a 10% higher win rate than humans gave them, and Claude-v1 with a 25% higher win rate. Then they declined the conclusion, writing that “due to limited data and small differences, our study cannot determine whether the models exhibit a self-enhancement bias” (Zheng et al., 2023). GPT-3.5 showed no preference for itself in the same comparison.
Those two figures travel widely without the sentence that follows them. They carry no interval, they come from a comparison the authors themselves read as underpowered, and the same figure shows the judges favoring models other than themselves. Self-preference bias is the name the later measurement work uses, and it claims less, describing a scoring gap without attributing a motive to the model.
Self-preferencing in competition law names something else
Regulators use self-preferencing for a platform ranking its own products above rivals’ inside a marketplace it operates, and Article 6(5) of the EU Digital Markets Act prohibits designated gatekeepers from treating their own services more favorably in ranking, indexing and crawling. The shape rhymes with the eval problem, an operator scoring a contest it has entered, and so does the remedy of putting distance between the scorer and the scored. Nothing measurable carries across: a competition case turns on conduct in a market, while the quantity on this page is a win-rate gap on a fixed set of pairs.
Neither self-preference nor verbosity can be read while the other stays uncontrolled, and both land on the close calls where comparing two answers rather than grading one decides a rank, so the two belong on one test plan. Every documented distortion has a detection test and a correction attached to it in LLM-as-a-judge bias, and the tests that catch it. To run those tests in a fixed order, each against an explicit pass line, work from the judge bias checklist.
Footnotes
-
Wei-Lin Chen, Zhepei Wei, Xinyu Zhu, Shi Feng and Yu Meng, “Do LLM Evaluators Prefer Themselves for a Reason?”, arXiv preprint. Scores self-preference against objective ground truth on mathematics, factual-knowledge and code benchmarks, and reports that stronger models show more pronounced harmful self-preference on the items they get wrong: https://arxiv.org/abs/2504.03846 (as of 2026-08). ↩
-
Yann Dubois, Balázs Galambosi, Percy Liang and Tatsunori B. Hashimoto, “Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators,” COLM 2024. Fits a generalized linear model on the auto-annotator’s preferences and predicts at zero length difference, raising Spearman correlation with LMSYS Chatbot Arena from 0.94 to 0.98: https://arxiv.org/abs/2404.04475 (as of 2026-08). ↩
-
Koki Wataoka, Tsubasa Takahashi and Ryokan Ri, “Self-Preference Bias in LLM-as-a-Judge,” NeurIPS 2024 Safe Generative AI Workshop. Introduces a quantitative metric for the bias and reports that LLM judges score low-perplexity outputs above the rating human evaluators give them, whether or not the text was self-generated: https://arxiv.org/abs/2410.21819 (as of 2026-08). ↩