Curbing the Winner’s Curse: Reward Aggregation Mitigates Inference-Time Reward Hacking
Abstract
Inference-time alignment methods such as Best-of- steer a language model with a reward model, but because every reward model is an imperfect proxy, optimizing it too hard backfires: true quality rises, peaks, and then collapses, a failure known as inference-time reward hacking. A recent line of work hedges the optimizer, calibrating how much to sample so as to stop at the peak. We ask whether this is the most effective lever available, and show that it is not. Separating how far to optimize from what to optimize exposes two label-free levers. The optimizer lever is nearly exhausted: the ideal amount of optimization is strongly per-prompt, so an oracle with labels would gain roughly points, but this signal is label-gated and our best label-free per-prompt selector recovers under half a point. The decisive lever is to fix the reward: given a suite of reward models, we aggregate their per-prompt-standardized scores by a low quantile across models rather than their mean, a value-at-risk (VaR) rule that better aligns the proxy and flattens the hacking curve. Across five verifiable domains, three inference-time mechanisms (BoN, SBoN, and BoP), and reward models, VaR aggregation raises expected accuracy by points over hedging a typical single model, an order of magnitude more than the optimizer lever yields. A mechanistic, falsifiable account explains when it helps: aggregation rejects a response favored by one model's blind spot precisely when reward-model errors are diverse, which predicts both its gains on reasoning and QA and its failure on code, where errors are correlated. It is also immediately deployable: training-free, with no added parameters, a two-line reduction over reward-model scores, and a small, diverse set of models recovers the full gain, keeping it compute- and memory-light. These results recast inference-time reward hacking as a reward-model deficiency to be repaired, not merely an optimization budget to be tuned.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.