acceptodds
Under review as a conference paper at ICLR 2027

Self-Preference Is a Sum: A Frame-Conditional Decomposition for LLM Judges

Abstract

Language models can encounter a prior answer alongside conflicting information. We test whether describing that answer as the model's own changes its choice. On held-out PopQA questions, we evaluate Qwen3, Qwen2.5, Phi-3.5 and OLMo-2. We pair a prepared answer with a conflicting document; either the answer or the document is correct. The answer text and number of appearances stay the same, verified in code; its description names the model itself, an unspecified language-model assistant or a different assistant. When the prior answer is wrong, self-attribution increases its selection by 7.5, 18.4 and 20.4 percentage points in three families relative to different-assistant attribution. This post-hoc analysis derives choices from answer-letter log probabilities. Different-assistant wrong-choice rates are 1.4%–6.8%; the fourth family's −0.7-point change is not statistically significant. Comparing own with unspecified authorship, and unspecified with different authorship, gives two log-probability differences that sum exactly to the total on the same questions. In a post-hoc analysis, rephrasing changes this total in all eight combinations of model family and prior-answer correctness. Only three remain after excluding conditions that choose one answer side at least 90% of the time, all with correct prior answers. We withdraw earlier interpretations that these differences vary independently or differ in their robustness to wording. Nine of fifteen retrospectively selected conditions miss their original criteria; none is established under our reporting policy adopted after seeing the data. The core study's preregistration remains unverified. Separately, relabelling model-generated answers from the model's own to an unspecified assistant's improves final-decision accuracy by 6.4–9.4 percentage points in three families. After exploratory grading exclusions, two gains retain confidence intervals above zero. Relabelling can also harm decisions on correct prior answers. Authorship wording therefore affects decisions, with outcomes depending on phrasing, analysis and prior-answer correctness.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.