PluralGen: Efficient Test-Time Scaling by Decoupling Particle Search from Model Execution
Abstract
Particle search gives compact language models a way to improve reasoning at test time when deployment constraints rule out a larger model. Resampling directs a fixed population toward promising paths but often leaves many particles at the same complete state. Sequence runtimes nevertheless decode every particle separately. Prefix-aware KV sharing preserves inherited history but repeats the current Transformer pass and vocabulary projection, so execution remains proportional to particle count even as the search concentrates. We present PluralGen, which keeps particles as the unit of search while advancing each distinct complete state once. It stores each repeated state with its multiplicity, draws an independent continuation for every particle, and carries the grouping through model execution, resampling, and KV ownership. PluralGen reaches 3.2× end-to-end speedup, reduces KV allocation by 88.1%, supports 4× larger populations, and improves serving throughput by 2.9× on the evaluated workloads.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.