Fewer Tokens, Adaptive Thinking: Dynamic Routing Framework with Soft budgets
Abstract
Chain-of-Thought (CoT) improves large language model reasoning but requires costly token-by-token decoding. Latent reasoning reduces explicit decoding but can be unstable and opaque. Hybrid reasoning combines sampled tokens with continuous context. However, using Hybrid inputs at every step may introduce noise and leave the token-efficiency potential of latent reasoning underused. To address these limitations, we propose DASH, a Dynamic routing framework with Adaptive Soft budgets for Hybrid reasoning. At each step, DASH selects between Token-only and Hybrid candidates using the current reasoning state and predictive uncertainty. This makes continuous-context use state-dependent while retaining token anchoring. An adaptive soft budget guides the reasoning scale without truncating generation. DASH learns routing and budget allocation through outcome-based reinforcement learning (RL). A three-stage curriculum establishes reliable routing, rewards shorter correct trajectories, and adapts the soft budget. No CoT annotations or intermediate labels are required. This avoids costly supervision and lets DASH discover reasoning paths that better balance correctness and length efficiency. Across five mathematical benchmarks and three model scales, DASH achieves higher accuracy with shorter trajectories than supervised fine-tuning with CoT (SFT-CoT) and outperforms representative latent-reasoning baselines. Compared with Hybrid Reasoning Policy Optimization (HRPO), it produces substantially shorter trajectories at an accuracy cost, achieving higher accuracy gain per reasoning position on four arithmetic benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.