acceptodds
Under review as a conference paper at ICLR 2027

Path-consistency with Prefix Enhancement for Efficient Inference in LLMs

Abstract

Self-consistency improves reasoning by aggregating multiple sampled solutions, but generating each trajectory from scratch incurs substantial inference cost. We propose path-consistency, an efficient test-time scaling method that reduces this cost through prefix enhancement, a controlled process that progressively extends a shared reasoning prefix across sampling windows and validates each extension with subsequent samples. Answer agreement guides candidate extensions, while trial and recovery determine whether to commit them or roll back to an earlier prefix. All valid answers from completed windows contribute to the final vote. We derive an exact finite-horizon decomposition of the gains and losses in voting accuracy induced by prefix reuse. We also establish a voting error bound for adaptive sampling with finitely many competing answers under explicit conditional-margin assumptions. Our evaluation targets three models of different scales, Llama-3.1-8B-Instruct, Qwen3-30B-A3B, and DeepSeek-V4-Flash, across five reasoning benchmarks. Across diverse models and reasoning benchmarks, path-consistency reduces generated tokens by 25.8% to 42.0%, averaging 34.4% across the long-chain settings, while largely maintaining accuracy. Combining path-consistency with early-stopping methods yields an additional reduction of approximately 20%. These results position path-consistency as a competitive reasoning-sharing method that also integrates flexibly with early-stopping methods for additional savings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.