General Preference Reinforcement Learning
Abstract
Online reinforcement learning (RL) with verifiable rewards drives emergent reasoning on math and code, but open-ended tasks have no programmatic verifier. A learned scalar reward model can be applied to these tasks in place of a verifier, yet a scalar is the wrong shape for the job. Quality is multi-dimensional, and any scalar score is an incomplete proxy that lets extended online RL collapse onto whichever axis the score is most sensitive to. Recent multi-reward methods counter this collapse by normalizing each reward separately, but they require every reward to be specified in advance. We instead learn latent quality dimensions from ordinary binary preference labels with the General Preference Model (GPM), and propose General Preference Reinforcement Learning (GPRL), which keeps these dimensions separate through the policy update. GPRL computes group-relative advantages per dimension, normalizes each on its own scale, and aggregates them with per-dimension weights. We characterize when this aggregate rejects responses that improve one dimension at the expense of the others, even when a scalar reward would favor them. The same structure yields a drift monitor that detects such exploitation during training and corrects it online. GPRL reaches length-controlled win rates of from and from on AlpacaEvalĀ 2.0 with the shortest responses of any method; it also outperforms offline, iterative, and scalar-reward online baselines on Arena-Hard, MT-Bench, and WildBench, and does not degrade under extended training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.