Scaling Feedback Alone Is Not Scaling Supervision: From Prediction Quality to Decision Quality in Reward Learning
Abstract
Scaling behavioral feedback data improves reward-model prediction but can reduce downstream decision quality. On an industrial dialogue agent serving over 100,000 users per group, increasing visit-event training data from 50% to 100% raised prediction AUC by 0.018 yet lowered order conversion by 0.15 percentage points (p<0.0001) through the full RM to RL to deployment pipeline. This prediction–decision mismatch extends to frozen-pool settings: across five conditions on tau^2-bench and VisualWebArena, better feedback fit accompanies worse candidate selection, with utility losses from -0.04 to -0.25. We propose a loss-preserving allocation diagnostic that exactly decomposes utility change into entry-mass and within-label ordering components, and derive an exact ceiling C-D on local repair gain within feedback-equivalent top sets. The ceiling holds universally—zero violations across 4,000+ task-configuration instances in three tau^2-bench domains—and significant positive repair appears under weak base selectors, confirming the diagnostic is actionable. A second online experiment with a random-down weight control isolates the mechanism: semantic filtering of ambiguous positive labels—not regularization from reduced volume—drives the improvement.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.