Adaptive Test-Time Compute Allocation for Reasoning LLMs via Constrained Policy Optimization
Abstract
Spending more compute at inference time can improve large language model (LLM) accuracy, but not every query benefits equally. Under a limited budget, which queries should receive more compute, and which can be answered with less? We formulate this decision as a constrained optimization problem: assign a compute budget to each query to maximize expected accuracy, subject to a limit on average compute cost. To solve this problem, we propose a two-stage Solve-then-Learn (STL) framework. In the solve stage, we use Lagrangian relaxation to construct a budget-allocation oracle with provable optimality guarantees on a calibration set. In the learn stage, we use the oracle allocations as supervision to train a lightweight classifier that predicts a compute budget for each new query before generation. For LLM reasoning, STL assigns different token caps to different questions to improve overall accuracy under a shared average token budget, while keeping the underlying LLM unchanged. Experiments across two benchmarks and three LLMs show that STL outperforms the strongest evaluated baseline in 17 of 18 settings, with an average accuracy gain of 3.90 percentage points and a maximum gain of 8.46 points under the same budget constraint.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.