acceptodds
Under review as a conference paper at ICLR 2027

PreSTO: Predictive Subtree Prefetching for Fast LLM Power Sampling

Abstract

Power sampling improves large language model (LLM) reasoning at test time by sampling from a sharpened distribution over the base model. Its sampler, PowerMH, is slow because its Metropolis–Hastings (MH) updates are sequential: each accept/reject decision sets the state for the next proposal. Prior work cuts this cost by changing the sampler or the model (e.g., sequential Monte Carlo, approximate lookahead, entropy-guided proposals, distillation). We instead keep the sampler: Predictive SubTree Prefetching (PreSTO) completes multiple MH transitions per model call without changing the MH transition kernel. Future MH outcomes form a tree of accept/reject decisions; PreSTO builds a prefetching subtree of it, generates every proposal whose prefix is already available in one batched call, and follows the realized path with the original MH acceptance rule. Because MH reaches only a sparse, nearly unary part of each prefetched tree, we choose which decisions to prefetch under a budget by dynamic programming. On PowerMH and its follow-ups EntropyCut and MultiTryMH across 7 datasets and 5 base LLMs, PreSTO achieves median speedups of , , and , respectively; on the first two, it makes a median of fewer model calls, each only longer. Our predictive traversal rule improves the mean speedup over the default breadth-first traversal.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.