Regime-Dependent Inference-Time Compute Allocation for Agentic Systems: When Deeper Reasoning Wins
Abstract
Longer reasoning, repeated sampling, revision, and communication have each been advocated as ways to improve language-model performance through their composition into agentic systems. These competing prescriptions leave a practical question: which use of compute is optimal for a given model, task, and compute budget? We organize this choice around the generation, preservation, and selection of correct answers, and compare 96 configurations of one model on four reasoning benchmarks. At a matched budget of 16,384 reasoning tokens, depth gives the highest observed accuracy in three harder groups defined retrospectively by short-run success. In the frequent-success group, voting reaches the observed ceiling at one-quarter of that expenditure. We then ask whether allocation can be prescribed before testing larger budgets on a new problem. A logistic rule trained on graded development outcomes at several depths chooses depth–width allocations using five short probes per new item. In an exploratory comparison on 200 unseen problems, it reaches 70.0% accuracy, versus 70.5% for fixed depth. The results support conditional allocation guidance and show how a calibrated rule can act before a new problem's scaling curve is measured.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.