Maximin Inference-Time Alignment with Diverse Rewards
Abstract
Inference-Time Alignment (ITA) improves a learned model’s output through sampling and search heuristics without modifying its underlying weights. While effective in guiding the model response, in reality a single model often serves a diverse group of users or is evaluated over multiple categories (e.g. accuracy, safety) with different reward models. This raises the question of how a shared and finite sampling budget should be allocated to ensure response quality across multiple groups of users or different evaluation criteria. We formulate this problem as a maximin budget allocation problem and study ITA under exact-reward and proxy reward settings. In the exact-reward setting, we upper bound the maximin regret of a greedy variant of Best-of-N algorithm (Greedy-BoN) and show that the regret goes to zero as the budget goes to infinity. However, in the proxy reward setting, the above approach fails even as the sampling budget goes to infinity. To overcome this, we develop Batched Greedy Pessimistic (BGP) algorithm that allocates each batch to the group using pessimistic response selection, and bound its maximin regret with respect to the proxy reward floor. Empirically, we evaluate Greedy-BoN and BGP on a dataset with five different domain (groups) and show that Greedy-BoN performs best with respect to the proxy reward, whereas BGP performs best when evaluated against an external judge, highlighting the importance of pessimism in inference-time alignment.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.