StepRouter: From Effort Priors to Utility Posteriors
Abstract
Chain-of-Thought (CoT) and related reasoning paradigms rely on historical tokens—intermediate conclusions from prior steps. Not all historical tokens are equally useful, and the lost in the middle phenomenon makes indiscriminate accumulation harmful. We study how to rank and select historical tokens, organizing them at the granularity of steps, which manifest as lemmas in mathematics, TODOs in code, or other domain-specific units. As a first step, we compare two ways to rank candidate steps. The first is an effort prior, which favors steps that look expensive to verify. The second is an ex-post utility score, which measures how much a step improves the final answer per token, judged by a frontier model. Empirically, effort is nearly uncorrelated with true usefulness, while utility-based ranking substantially outperforms effort and other baselines at selecting high-value steps. Despite the simplicity of these signals, we develop StepRouter, a lightweight posterior utility estimator that guides any backbone to rank steps by utility, routing only the top 20% for reuse. The results are surprisingly strong: StepRouter improves all tested models on several reasoning benchmarks, yielding +19.6 pp on LiveCodeBench with DeepSeek-R1, +23.2 pp with Qwen3-8B, and pushing GPT-5 to 97.2% on AIME. Fine-tuning Qwen3-8B with only 166 examples yields +12.2 pp on AIME (68.16%→80.36%), demonstrating that models can internalize effective token utilization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.