acceptodds
Under review as a conference paper at ICLR 2027

Catching RLHF Failures before Training

Abstract

RLHF plays a key role in post-training, but the most reliable way to debug an RLHF setup, including its reward model (RM), is still running the training end to end: static RM benchmarks and Best-of-N (BoN) sampling correlate poorly with downstream outcomes. We introduce Lookahead, a sampling strategy that anticipates training failures before they occur, using only the base model and the RM. At selected decoding positions, each candidate token's logits are shifted by the reward that a rollout continuing from it obtains, with a coefficient controlling the optimisation pressure applied. At high , the RM's preferences are amplified until its flaws become visible. Across four (base model, RM) pairs spanning two model families and two scales, Lookahead outperforms BoN at predicting which IFEval prompts degrade after training, by a margin of 0.11–0.26. Increasing the BoN sample budget up to does not close this gap. At the instruction-family level, our method flags 20 of the 24 degradation cases, while BoN wrongly predicts improvement in 17. Our results show that much of what RLHF will break is already determined by the (base model, RM) pair, and can be read out before training begins.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.