A Moving Ruler: Evaluation Comparability Across RLVR Checkpoints
Abstract
Checkpoint evaluations are routinely used to quantify the evolution of model behavior with reinforcement learning from verifiable rewards (RLVR). Naturally, a robust evaluation demands a consistent scoring protocol at every checkpoint, which we refer to as evaluation comparability. However, evaluation comparability across RLVR checkpoints is rarely analyzed, even though RLVR training changes the nature of the response that the evaluation receives. For example, in AllenAI's Olmo3.1 RL-Zero Math model, benchmark scores rise over training checkpoints. But we observe that most of the gain is a result of the model's ability to deliver an answer within the token budget, rather than an increased capability in solving problems. Furthermore, we observe that later checkpoints also stall less often, and the probability of repeating the final answer within the chain-of-thought (CoT) increases. These changes in the nature of the response do not remain within the model, and they leak into the evaluation itself. To address this problem, we probe the model with a panel of problems with poisoned answer keys under various interventions. We identify three channels through which an evaluation harness that uses an LLM judge reacts to these changes. Each channel shifts the judge labels, which shows that the behavioral curve across checkpoints is altered without any change in the behavior the judge claims to measure. We finally distill these observations into a four-practice protocol. Our work shows that researchers must ensure evaluation comparability before reading a trend across RLVR checkpoints as a change in capability or safety.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.