BoNVoyage: Learning Better Rewards without Ranking
Abstract
Reward models (RMs) play a central role in reinforcement learning from human feedback (RLHF), yet they are typically trained with a Bradley-Terry (BT) objective that ranks response pairs under a fixed text distribution. This objective does not match how RMs are used in RLHF: as training progresses, the policy increasingly concentrates on responses that receive high reward under the RM itself. But what matters is not whether the RM ranks fixed pairs correctly, but whether it remains reliable on the responses that its own signal makes more likely under the policy. We develop BoNVoyage, a method for training RMs directly for this downstream role. Instead of training a pairwise ranker, BoNVoyage maximizes the likelihood that the RM-induced RLHF-optimal policy assigns to preferred responses. Estimating the gradient of this objective requires samples from the optimal policy induced by the learning RM. We obtain these samples through test-time alignment: given the base LM and the RM, we use Markov chain Monte Carlo to sample from the induced policy. To make this practical, we use the base LM as an independent proposal distribution, reuse candidate responses across optimization steps, and use contrastive divergence with short chains initialized at preferred responses. We evaluate BoNVoyage by training 4B RMs for Qwen3-1.7B-Base in mathematics and our supervised finetuned model Qwen3-1.7B-Science in science. Compared with BT baselines, BoNVoyage yields stronger downstream policies under RLHF and stronger rerankers under best-of- sampling across MATH-500, GSM8K, SciQ, SciKnowEval, ARC-Challenge, and MMLU-Science. On verifiable math tasks, BoNVoyage also shows substantially more robustness to reward over-optimization, preserving alignment between RM scores and ground-truth correctness later into RLHF training. On publication, we will release our code, trained models, and datasets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.