acceptodds
Under review as a conference paper at ICLR 2027

SiftKV: Verifier State Reuse across Search Rounds for Efficient Test-Time Scaling

Abstract

Verifier-guided test-time scaling improves reasoning by expanding several candidate paths and retaining the most promising ones for further search. The selected paths return as the starting points of the next round, so the same reasoning histories continue across rounds. However, conventional process reward model (PRM) serving scores every extension as a new, complete request. It repeatedly reconstructs the histories of surviving paths and materializes KV for candidates that selection immediately removes. As the tree grows, an increasing share of verifier work repeats past computation instead of advancing the search. We introduce SiftKV, a PRM serving system that carries verifier state with the selected paths across rounds. Each surviving path resumes from a token-consistent continuation that preserves its causal KV and accumulated step rewards. New suffix state remains provisional until selection, after which only the selected suffixes join a tree that stores shared ancestry once. Verification can then extend the surviving paths without reconstructing their histories. Across reasoning workloads, search policies, and generator–verifier pairs, SiftKV delivers up to 50% lower end-to-end latency and 69% lower verifier time, together with up to 104% higher Precise Goodput.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.