acceptodds
Under review as a conference paper at ICLR 2027

Adaptive Utility-Rational Allocation: Thinking Only When It Matters

Abstract

Reasoning models can continue generating after an answer has stabilized, increasing inference cost without improving task performance. We introduce Adaptive Utility-Rational Allocation (AURA), a controller framework that compares a proxy for further improvement with the price of additional reasoning tokens. AURA combines a development-tuned certainty score, stopping safeguards, bounded auxiliary checks, and batch-level price feedback while leaving model weights fixed. On MATH-500, AURA achieves 93.2% accuracy with 646 tokens, versus 748 for Plan-and-Budget and 751 for an RL halter at the same accuracy point estimate. Paired intervals support these 13.6–14.0% token reductions. Score ablations on MATH-500 and TeleQuAD identify confidence as a consistent contributor to token efficiency, while a confidence-only score nearly reproduces the full MATH result and uses fewer tokens on TeleQuAD. In cumulative ablations across both tasks, safeguards and auxiliary checks are associated with higher observed answer quality. Batch-level price feedback reduces the remaining budget overshoot. We relate utility-based stopping to certainty thresholds and show when estimation errors leave the local stopping decision unchanged. The results support a modular view of adaptive stopping: score composition, eligibility safeguards, and budget feedback govern different aspects of the accuracy–cost tradeoff.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.