When Is More Reasoning Worth It? Fallback-Aware Reasoning-Strategy Routing Under Hard Token Budgets
Abstract
Reasoning strategies have heterogeneous, input-dependent token costs. Standard test-time allocation asks where compute is useful; we study the distinct problem of hard per-query output-token compliance, where realized use above a requested threshold is a violation. We separate this risk into violations when a strategy is predicted feasible and violations after no strategy is predicted feasible, then use a frozen router combining strategy-specific token prediction, estimator-fit accuracy priors, conservative resource corrections, and minimum-predicted-token fallback. In a sealed true-online evaluation of Qwen2.5-3B-Instruct and Phi-3.5-mini-instruct across four reasoning benchmarks, 2,400 routed executions precede a 4,000-execution exhaustive matrix. From low to high budgets, macro violation falls from 0.258 to 0.033 for Qwen and from 0.225 to 0.038 for Phi, while fallback falls from 0.688 to 0.123 and from 0.750 to 0.048, respectively. Yet conservatism is not uniformly beneficial: against the frozen point router, Phi’s low-budget MATH-500 violation increases by 0.25 (95% bootstrap CI [0.15, 0.35]; Holm-adjusted p = 0.00024), and low-budget fallback carries substantial model-dependent accuracy costs. Hard-budget reasoning therefore requires controlling both uncertainty within the feasible branch and the behavior of the system when predicted feasibility disappears.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.