When Position Predicts Reward in LLM Judges
Abstract
LLM judges can learn a position rule instead of comparing response quality when preference data always place the preferred response in the same slot. We study this training-interface confound as an identification problem: a content-sensitive judge should reverse its verdict when the two responses are exchanged, whereas a position shortcut should not. We introduce a paired-order audit and train matched judges with either a fixed preferred-first interface or both physical arrangements. In a controlled four-arm study with Qwen2.5-7B, four seeds, and 200 update steps, two-order pointwise training raises order-averaged accuracy from 65.14% to 87.05% on RewardBench and from 55.51% to 75.31% on an UltraFeedback clean split. First-position selection falls from 83.30% to 49.73% and from 93.12% to 56.18%, respectively. A separate dose study shows that increasing the preferred-first share monotonically reduces paired robustness. The results establish a simple audit and training intervention for testing whether an apparent judge gain survives a change to the feature that predicts its reward.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.