Learning From Everything: Dense Rewards From Every Kind of Feedback
Abstract
Learning jointly from different kinds of feedback remains an open challenge in language model post-training. Beyond verifiable and rubric-based rewards, current pipelines rely primarily on demonstrations and preferences, leaving other useful signals underused. We propose a unified and scalable method for learning dense rewards jointly from demonstrations, preferences, edits, ratings, stops, and more. We model each feedback type through its own likelihood and use Bayesian inference to learn a shared distribution over token-level rewards. The resulting probabilistic reward model provides dense rewards with uncertainty estimates and accommodates any feedback that can be modeled through a likelihood over rewards or action values. Training a model from diverse feedback then reduces to learning one shared reward model and using its rewards to optimize the policy. On the mathematical reasoning and safety subsets of RewardBench, the best feedback combinations improve accuracy over preferences alone by up to 17.4 and 5.9 percentage points, respectively, and over standard sequence-level reward models by 19.8 and 17.8 points. The learned rewards also improve error localization within responses and support stronger, more sustained improvements during PPO on mathematical reasoning tasks. Together, these results demonstrate the value of a unified reward model for turning diverse feedback into better policies.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.