WHEN SELF-PREFERENCE IS A VERIFICATION FAILURE: CAUSAL EVIDENCE FROM VISION-LANGUAGE JUDGES
Abstract
Vision-language models used as judges tend to prefer their own answers, and the usual explanation is narcissism: the judge recognizes its own output and rewards it. We argue instead that self-preference is mostly a verification failure: a judge checks each answer against its own interpretation of the image, so when it misinterprets the image, it confirms the wrong answer produced from that misunderstanding. We test this causally in 19 open vision-language judges from 8 model families, across five multimodal benchmarks and a text control, by changing the judge's image while keeping every candidate answer byte-identical. An outcome-matched control shows that most raw self-preference is evaluator noise, consistent with recent findings in text-only settings. The remaining bias follows the judge's interpretation of the image rather than authorship: answers are favored for the scene they were written from rather than because the judge wrote them. Degrading the judge's view steadily reduces this bias, and showing the judge the image used to generate its other answer reverses this preference. A judge's accuracy at verifying other models' answers on items it answers incorrectly predicts its favoritism, including for judges and families excluded from fitting; the signal comes from answers that repeat the judge's own mistake. Tests on nine larger judges, four commercial, show that this relationship predicts their much higher favoritism, but an authorship effect can no longer be ruled out: individual larger judges show substantial direct authorship preference, the strongest almost matching that judge's evidence effect. This mechanism has a practical cost and a diagnostic use: a model choosing among its own samples selects a correct answer less often than other judges do on average, and removing or replacing the image during evaluation changes which judges are identified as biased.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.