acceptodds
Under review as a conference paper at ICLR 2027

Scaling Feedback Alone Is Not Scaling Supervision: From Prediction Quality to Decision Quality in Reward Learning

Abstract

Scaling behavioral feedback data improves reward-model prediction but can reduce downstream decision quality. On an industrial dialogue agent serving over 100,000 users per group, increasing visit-event training data from 50% to 100% raised prediction AUC by 0.018 yet lowered order conversion by 0.15 percentage points (p<0.0001) through the full RM to RL to deployment pipeline. This prediction–decision mismatch extends to frozen-pool settings: across five conditions on tau^2-bench and VisualWebArena, better feedback fit accompanies worse candidate selection, with utility losses from -0.04 to -0.25. We propose a loss-preserving allocation diagnostic that exactly decomposes utility change into entry-mass and within-label ordering components, and derive an exact ceiling C-D on local repair gain within feedback-equivalent top sets. The ceiling holds universally—zero violations across 4,000+ task-configuration instances in three tau^2-bench domains—and significant positive repair appears under weak base selectors, confirming the diagnostic is actionable. A second online experiment with a random-down weight control isolates the mechanism: semantic filtering of ambiguous positive labels—not regularization from reduced volume—drives the improvement.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.