When Do LLM Judges Favor Themselves? Dissociating Self-Recognition from Self-Preference
Abstract
LLM-as-a-judge has become a standard evaluation paradigm, yet language models can exhibit self-preference, favoring their own outputs over those of other models. Existing explanations attribute this effect to self-recognition, familiarity, or genuine quality differences, but these factors have rarely been isolated within a unified controlled setting. We re-examine self-preference under stricter controls by fixing candidate pairs to identical incorrect answers and clustering inference at the question level. Across seven judges and four tasks, we find that self-preference persists, but is substantially smaller and more task-dependent than conventional pairwise analyses suggest. The effect increases with model scale and is strongest on tasks with greater expressive latitude, reaching +0.24 on MMLU, whereas it vanishes or reverses on verifiable numerical tasks. Our linear-probe experiments recover self-authorship with near-perfect accuracy across models from 1B to 72B parameters, even after surface cues are neutralized. Crucially, however, self-recognition and familiarity remain strong even when self-preference disappears, showing that neither is sufficient to explain the bias. After controlling for judge capability, the residual self-specific effect is only about two percentage points, yet excluding self-votes can still change the top-ranked model. These results show that self-preference is real but modest, strongly task-dependent, not reducible to surface features, and potentially consequential for model evaluation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.