Think Before Spending: Value-Aware Computation Control for LLM Reasoning
Abstract
Test-time computation improves large language model (LLM) reasoning by generating additional evidence and exploring multiple reasoning paths, but incurs substantial generation cost. Existing adaptive computation methods mainly optimize resource allocation or termination after a reasoning strategy has been fixed, leaving the initial strategy choice unresolved. Reasoning strategies such as direct submission, evidence gathering, and consistency sampling offer different quality–cost trade-offs across requests, making the choice of strategy central to reducing unnecessary computation. To address this challenge, we present Compass, a value-aware computation controller for LLM reasoning. Compass is built on two key observations: (1) observable inference states contain signals about the expected quality and cost of candidate reasoning strategies, and (2) additional computation is valuable only when it improves upon the currently available answer. Based on these observations, Compass introduces two coupled decisions: quality-aware computation strategy selection and marginal-value continuation control. Specifically, Compass estimates the final quality and remaining token cost of candidate computation strategies before committing to expensive execution, selecting a low-cost strategy with near-optimal predicted quality. During execution, Compass dynamically reassesses continuation by estimating the probability of correcting the current answer against the risk of introducing an error and the additional computation cost. Across four datasets, three starting models, and two generation seeds, Compass reduces pooled output tokens per correct answer (TPC) by 20.1–88.1% relative to the evaluated sampling baselines, with macro-accuracy gains of 0.9–2.3 percentage points.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.