SMART: Systematic Measurement and Adaptive Routing of Thinking for LLM Agents
Abstract
How much should an LLM agent think before it acts? We present SMART (Systematic Measurement and Adaptive Routing of Thinking), a study that treats the per-step thinking budget—the maximum number of reasoning tokens allowed per step—as an economic quantity to be measured and optimally spent. Across 11,952 evaluation episodes on three agentic domains, budget sensitivity proves environment-dependent: success–budget curves saturate quickly in short-horizon dialogue but keep scaling in long-horizon code execution. We trace the pathology of mid-range budgets to a truncation tax: reasoning cut off mid-thought degrades into garbled steps that pollute the context. The pollution inflates later steps' token use by 13.2% on average (9–31% across domains), making intermediate-high budgets the least favorable trade on the grid. A 1.7B router trained on successful trajectories is Pareto-optimal in every domain—beating the pre-specified compute-optimal budget at 48.5% (airline) and 26.5% (retail) fewer reasoning tokens, and reaching the Pareto frontier in AppWorld, where success keeps scaling with budget—and the optimal allocation granularity matches the domain's interaction horizon: per-step in dialogue, per-task in long-horizon execution. Finally, we examine what paid-for thinking contains: budget vigorously reshapes what agents think about, yet no behavior predicts success positively—backtracking instead marks struggle, not failure. What budget buys is not a different mix of thoughts but intact ones.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.