Beyond Fixed Rules: Best-of-N Alignment with Learned Set-wise Selection
Abstract
Best-of-, which generates candidate responses from an LLM and selects the one with the highest reward-model score, is widely used for inference-time alignment but is susceptible to reward hacking as grows. Some recent work seeks to mitigate this problem by designing proxy-reward correction rules based on signals over the candidate responses. However, these methods encode predefined assumptions about reward-model errors and the rigidity of such fixed-form rules may limit their ability to capture reward-model error patterns that vary across inference-time conditions, making it difficult to reliably distinguish reward-hacked responses from genuinely high-quality ones. To move beyond fixed rules under varying inference-time conditions, we propose LESS, a LEarned Set-wise Selection framework for Best-of- alignment, which learns a tailored selection strategy from the corresponding data to select the best response. It represents each candidate using reward and representation features, then employs a set transformer encoder to capture interactions among the candidates and assign each response a selection logit, returning the highest-scoring candidate as the final output. The set-based design naturally accommodates varying candidate-set sizes. Extensive experiments across diverse benchmarks and varying values of demonstrate the effectiveness of the proposed approach.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.