Rethinking Surrogate Criteria for RLVR Rollout Filtering
Abstract
Reinforcement learning with verifiable rewards improves language-model reasoning, but it remains unclear whether common criteria for filtering generated responses preserve the learning signal needed for policy updates. We investigate when directional agreement preserves the gradient computed from all responses and when magnitude mismatch changes which subset should be preferred. We develop a controlled audit that separates gradient error into direction and scale, compares all subsets at the same retention budget, and evaluates selected subsets in the original trainable parameter space. Our theoretical analysis shows that perfect directional alignment can still yield large reconstruction error, motivating a selection criterion that jointly matches direction and magnitude. Across eight model–dataset pairs covering GSM8K, MATH, and TACO, retaining eight of sixteen responses with our criterion reduces mean squared relative gradient error by 66.8–75.7% compared with reward-variance selection and 42.9–53.7% compared with direction-only selection. These results identify scale mismatch as a limitation of common selection criteria and provide a framework for evaluating gradient preservation at a fixed policy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.