acceptodds
Under review as a conference paper at ICLR 2027

Not All Turns Are Equally Hard: Adaptive Thinking Budgets For Efficient Multi-Turn Reasoning in Agents

Abstract

As LLM reasoning performance plateaus, improving inference-time compute efficiency is crucial to mitigate overthinking and long thinking traces even for simple queries. Prior approaches including length regularization, adaptive routing and difficulty-based budget allocation primarily focus on single-turn settings and fail to address the sequential dependencies inherent in multi-turn reasoning. We formulate multi-turn reasoning as a sequential compute allocation problem and model it as a multi-objective Markov Decision Process. We propose TAB: Turn-Adaptive Budgets, a budget allocation policy that takes as input the conversation history and learns to maximize task accuracy while respecting global per-problem token constraints, and consequently learns to adaptively allot smaller budgets to easier turns and save an appropriate number of tokens for the crucial harder reasoning steps. We train and analyze the behavior of TAB on an auxiliary hierarchical plan-and-execute system for multi-turn math reasoning where solution trajectories are inexpensive to collect. Here, TAB saves up to 35% tokens and 30% latency over static and off-the-shelf LLM budget baselines. We then transfer the same TAB policy to environment-grounded agentic workflows spanning sequential tool-use, simulated user-agent dialogue and multi-step terminal tasks, where it generalizes and maintains accuracy while saving up to 20% tokens and 15% latency, without ever being trained on them.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.