SABER: Stage-Adaptive Bridging and Early Routing for Large Reasoning Models
Abstract
Large reasoning models (LRMs) achieve strong performance through extended chains of thought, but their long reasoning trajectories incur substantial inference overhead. Existing early-exit methods reduce this cost by terminating reasoning once a correct answer becomes extractable, implicitly equating answerability with response readiness. In this work, we find that LRMs frequently become answerable well before they are ready to produce a high-quality response, revealing a Solution-Response Gap between answerability and response readiness, during which further computation consolidates the reasoning process and prepares the final response. To better exploit the information encoded within this gap, we introduce SABER, a stage-adaptive framework designed to harness the intermediate reasoning signals that bridge answerability and response readiness. SABER realizes this objective through two complementary modules: SCOUT identifies when the reasoning state becomes answerable, while BRIDGE efficiently exploits the remaining consolidation through compact recurrent latent computation. Across multiple reasoning benchmarks and different LRM backbones, SABER reduces average token usage by 22.1–49.6% while consistently outperforming competing compression methods in accuracy and quality. Beyond these performance gains, further analyses reveal heterogeneous consolidation patterns associated with distinct reasoning behaviors and varying contributions to response quality. These findings provide the research community with a functional perspective on reasoning consolidation, shifting the focus beyond when to stop toward what computation remains useful in achieving response readiness.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.