acceptodds
Under review as a conference paper at ICLR 2027

BranchWeave: Local Wavefronts with Device-Wide Batching for Verifier-Guided Test-Time Scaling

Abstract

Test-time scaling lets compact language models spend additional inference compute on difficult queries, extending reasoning under deployment constraints. Verifier-guided search converts this budget into candidate paths and scores, yet existing runtimes advance subtrees in global depth rounds. A completed subtree waits for unrelated work, while immediate dispatch produces small generator and verifier batches. We introduce BranchWeave, which separates the dependency scope of search from the batching scope of model execution. BranchWeave derives selection domains from reducer dependencies; once a domain receives its selection scores, its local wavefront advances. Stable node identities keep sampling decisions, scores, KV state, and speculative continuations attached to their owners as calls are reordered. Device-wide generator and verifier queues then pool ready calls across domains and depths, forming dense batches without delaying local progress. Together, local release and shared batching convert barrier time into scored reasoning before the response deadline. Across three generator–PRM pairs and five reasoning benchmarks, BranchWeave completes searches up to 5.46× faster and delivers up to 4.94× higher goodput.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.