Do LLM Evaluators Prefer Themselves for a Reason?
Abstract
Large language models (LLMs) are increasingly used as automatic evaluators in applications such as benchmarking, reward modeling, and self-refinement. Prior work suggests that these LLM evaluators exhibit *self-preference*: they favor their own outputs, often more strongly as model capability increases. However, whether self-preference behavior is really harmful, or simply reflects the genuinely higher-quality outputs of stronger models, has been overlooked. To resolve this, we study self-preference on verifiable benchmarks spanning mathematical reasoning, factual knowledge, and code generation. This setting enables us to separate *legitimate* self-preference (favoring objectively better responses) from *harmful* self-preference (favoring worse ones), different from past studies that focus on subjective tasks, where preferences are hard to verify objectively. Across large-scale experiments on Llama, Qwen, Gemma, Mistral, Phi, GPT, and DeepSeek models, we find three main results: *(1)* Stronger models show greater self-preference, but much of it is legitimate and tracks their genuinely better generations. *(2)* Harmful self-preference persists when models make mistakes as generators, and it becomes more pronounced in stronger models, suggesting they struggle more to recognize when they are wrong. *(3)* Inference-time scaling strategies, such as generating a long Chain-of-Thought before evaluation, effectively reduce harmful self-preference. Notably, we conduct further experiments with LMArena and observe similar patterns, showing our findings extend beyond verifiable tasks to subjective, real-world settings. Overall, our results offer a sharper picture of when self-preference is benign versus problematic, and point to practical ways to make LLM-based evaluation more reliable.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.