BranchWeave: Local Wavefronts with Device-Wide Batching for Verifier-Guided Test-Time Scaling
Abstract
Test-time scaling lets compact language models spend additional inference compute on difficult queries, extending reasoning under deployment constraints. Verifier-guided search converts this budget into candidate paths and scores, yet existing runtimes advance subtrees in global depth rounds. A completed subtree waits for unrelated work, while immediate dispatch produces small generator and verifier batches. We introduce BranchWeave, which separates the dependency scope of search from the batching scope of model execution. BranchWeave derives selection domains from reducer dependencies; once a domain receives its selection scores, its local wavefront advances. Stable node identities keep sampling decisions, scores, KV state, and speculative continuations attached to their owners as calls are reordered. Device-wide generator and verifier queues then pool ready calls across domains and depths, forming dense batches without delaying local progress. Together, local release and shared batching convert barrier time into scored reasoning before the response deadline. Across three generator–PRM pairs and five reasoning benchmarks, BranchWeave completes searches up to 5.46× faster and delivers up to 4.94× higher goodput.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.