Controlled Audits of Policy–Judge Coupling
Abstract
Self-rewarding language models reuse a policy as the judge supplying its next training signal. We distinguish changes in evaluator preferences from changes in evaluated responses through two controlled experiments on factual multiple-choice questions. First, policy fine-tuning varies the association between a user's stated answer (the stance) and the response while matching prompt and response marginals, without judge-format examples. On fixed responses, the judge's relative preference for agreement follows agreement-paired neutral disagreement-paired training in all nine model–seed combinations across Qwen3.5-9B, Qwen3.5-4B, and Llama-3.1-8B-Instruct. Every trained condition nevertheless has a lower stance-dependent score contrast than base. Second, at an identical policy checkpoint and candidate set, replacing the learned stance-sensitive reward component with its base-model counterpart changes 33.40–41.60% of selected preference pairs across three Qwen3.5-9B seeds. The next-update stance-dependence contrast is +3.55 percentage points, with mixed seed signs and a 95% interval of [−0.02, 6.23]. Averaging rewards over both stances for four rounds yields 10.75 percentage points lower measured dependence; capability preservation is unestablished, and unresolved answer readings permit either sign. Matched exposure reveals evaluator differences that base-only audits miss, while reward interventions require separate validation of selected supervision and policy behavior.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.