acceptodds
Under review as a conference paper at ICLR 2027

Annealed Predictor-Corrector Sampling for Inference-Time Scaling of LLMs

Abstract

Autoregressive large language models (LLMs) have demonstrated strong performance on various reasoning tasks, and when a reward model is available, inference-time scaling can steer them further toward high-reward responses. Best-of-N sampling draws responses independently from the base model and selects the highest-reward response, but reward information from one candidate is not used to guide the generation of others. Recent Sequential Monte Carlo (SMC) approaches instead guide generation along the token axis, requiring intermediate estimates of partial-response quality that may prematurely discard promising trajectories. We propose the Annealed Predictor-Corrector (APC) sampler, which performs SMC along the reward-tilting axis by constructing a sequence of intermediate distributions with gradually increasing reward strength. APC maintains complete responses as particles and alternates between predictor and corrector steps. The predictor progressively favors higher-reward responses through reweighting, while the corrector refines particles by regenerating response suffixes from the base model and selecting among candidates using a transition that preserves the current intermediate target distribution. Applying this correction along the annealing path enables responses to be progressively refined before the target distribution becomes strongly concentrated. We further introduce an adaptive stopping rule that avoids unnecessary computation once a response stabilizes. Across reasoning benchmarks, APC provides a favorable accuracy-compute trade-off compared with independent sampling and token-axis inference-time scaling baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.