Catching RLHF Failures before Training
Abstract
RLHF plays a key role in post-training, but the most reliable way to debug an RLHF setup, including its reward model (RM), is still running the training end to end: static RM benchmarks and Best-of-N (BoN) sampling correlate poorly with downstream outcomes. We introduce Lookahead, a sampling strategy that anticipates training failures before they occur, using only the base model and the RM. At selected decoding positions, each candidate token's logits are shifted by the reward that a rollout continuing from it obtains, with a coefficient controlling the optimisation pressure applied. At high , the RM's preferences are amplified until its flaws become visible. Across four (base model, RM) pairs spanning two model families and two scales, Lookahead outperforms BoN at predicting which IFEval prompts degrade after training, by a margin of 0.11–0.26. Increasing the BoN sample budget up to does not close this gap. At the instruction-family level, our method flags 20 of the 24 degradation cases, while BoN wrongly predicts improvement in 17. Our results show that much of what RLHF will break is already determined by the (base model, RM) pair, and can be read out before training begins.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.