ARC: Trajectory-Aware Adaptive Reasoning Control via Latent-State Dynamics
Abstract
Large Reasoning Models (LRMs) have demonstrated strong capabilities in solving complex tasks through extended chain-of-thought reasoning. However, existing approaches to efficient reasoning primarily rely on truncation, compression, or local step-level signals, limiting adaptive computation allocation across the reasoning trajectory. Our empirical analysis shows that local signals can be ambiguous across different reasoning stages, whereas hidden states capture richer information about the evolution of reasoning. Building on this observation, we propose Adaptive Reasoning Controller (ARC), a lightweight trajectory-aware controller that leverages the hidden-state history of a frozen reasoning model to select among four discrete actions—Deep, Normal, Compress, or Conclude—for regulating subsequent computation. ARC introduces fewer than 2.8M trainable parameters, a negligible overhead over the frozen reasoning model. ARC is initialized through supervised learning on latent reasoning trajectories using pseudo-action labels derived from PRM signals, and is further optimized with Group Relative Policy Optimization (GRPO) using final-answer correctness and reasoning cost as learning signals. Under this hierarchical formulation, the frozen reasoning model serves as the low-level reasoning policy, whereas ARC learns a high-level strategy for adaptive computation allocation. Experiments across 4 LRMs and 4 benchmarks show that ARC reduces reasoning tokens by 32.8% on average and up to 55.3%, while increasing macro-averaged accuracy from 58.43% to 58.79%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.