acceptodds
Under review as a conference paper at ICLR 2027

REWARDING THE WRONG CHARACTER: EVALUATOR VALIDITY UNDER POLICY OPTIMIZATION

Abstract

Reward models are used both to evaluate outputs and to optimize the policies that generate them. This dual role creates a validation problem: strong static discrimination may coexist with misleading assessments of policy progress. We investigate this problem in role-playing, where character authenticity can be assessed independently by human annotators. RoleRM achieves 95.3% accuracy on static character contrasts, yet optimizing Qwen3-8B against its fixed scoring rule increases held-out reward from 13.62 to 17.72 out of 18 while human authenticity falls from 3.54 to 1.47 out of 5. On 50 matched prompts, RoleRM favors the final response on 43, whereas humans favor the initial response on 47. We term this pattern reward–fidelity inversion. Reward saturation, loss of within-group score variation, and response homogenization accompany the divergence; stronger KL regularization attenuates it, and a second policy–reward configuration exhibits the same direction of change. We organize these measurements into a trajectory-based evaluator validation protocol that pairs fixed prompts and checkpoint comparisons with independent outcome judgments. The central lesson is that holding out prompts does not make the reward used for optimization an independent measure of progress.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.