acceptodds
Under review as a conference paper at ICLR 2027

What Does Static Reward-Model Evaluation Identify About Downstream Optimization?

Abstract

In reinforcement learning from human feedback (RLHF), static evaluations are commonly used to select reward models (RM) before policy optimization. We ask when such evaluations are sufficient to determine which RM yields a better policy. The question matters because training a policy for every candidate is expensive, while RMs with similar static accuracy are known to produce very different policies. Two RMs can rank every response identically and have the same reward variance, yet one improves the policy while the other makes it worse. Ranking and variance indicate how strongly an RM prefers some responses over others, not whether those preferences agree with the target's. The question is therefore not how many scores to observe, but whether they are taken on responses that the target itself distinguishes. We formalize this for reward models that are linear in a fixed set of features. Under a norm-bounded linear reward class, we derive the sharp range of initial effects consistent with exact score observations, showing that the pairwise comparison can be identified without recovering every reward score. This local comparison need not persist as the policy distribution changes. We derive sufficient conditions under which it remains identified after a finite policy change, including at matched KL distance. Experiments with pretrained RM scores show that partial observations can identify the comparison after finite optimization or leave both signs attainable. Static evaluation can thus guide RM selection, but only when it observes what the target cares about, and only for limited optimization.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.