Sequential Anytime Valid Tests for Efficient Inference-Time Alignment
Abstract
Inference-time alignment methods, such as Best-of-N (BoN), improve language model outputs by generating many candidates per prompt and selecting among them using a reward model. The cost is largely wasted, since a fixed budget is applied uniformly across prompts, even though the computation required by different prompts varies widely. Adaptive methods halt generation once the sampled answers appear settled, but their thresholds are tuned heuristics with no control over the probability of stopping at the wrong answer. Instead, we ask for a rule that takes that probability – a permitted error level – as input, and returns the number of tokens spent to meet it. To build such a rule, we aggregate candidates by their final answers using a reward-weighted self-consistency (RWSC) score and treat the decision of when to stop as a sequential test over the resulting clusters. Wald's sequential probability ratio test (SPRT) is the natural tool for such a test, and when applied to the RWSC clusters, it yields SPRT-RWSC, which performs well empirically. Yet its guarantee does not hold here, since the leading answer shifts as sampling proceeds, and the run is continuously monitored. Our main contribution replaces Wald's test with an anytime-valid construction that resolves this issue: it maintains a betting supermartingale (an e-process) for every answer the policy produces, bounding the probability of ever certifying a non-dominant answer by , uniformly over time and across all answers, however many the policy produces. We call the resulting rule e-RWSC: the RWSC scores drive the bets, which affect only how quickly a dominant answer is certified, not the validity of the guarantee; and certification takes samples. Experiments across several language models, datasets, and reward models confirm the predicted scaling of per-response cost, and show that accuracy improves over baselines at substantially lower token cost.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.