acceptodds
Under review as a conference paper at ICLR 2027

The Conditional Value of Compute: When Is Adaptive Inference Worth It?

Abstract

More test-time compute can improve LLM performance, but a fixed budget makes this benefit selective: which inputs should receive the extra computation? Average quality-cost curves do not answer this, and difficulty or confidence can be misleading because inputs that are hard for a cheap action may also be hard for a more expensive one. We therefore study the expected improvement from additional computation conditioned on information available before the decision-the conditional value of compute. We derive the optimal allocation rule across budgets, characterize what routing information must preserve to support these decisions, and quantify losses from estimation error and from paying to acquire the routing signal. On four SWE-Gym judge pairs, learned routing improves held-out trajectory-judgment accuracy by 3.92-8.35 percentage points over random allocation when 30% of inputs receive the more expensive judge, with positive gains in all 44 leave-one-repository-out evaluations. On RouterBench, question-text routing improves actual answer correctness by 3.32-6.02 points on three representative model pairs; 1.37-2.64-point gains remain when the upgrade budget is fixed separately within each task. Complementary experiments reinforce a second distinction: a signal can improve targeting yet may not be worth acquiring once its own cost and competing cheap actions are counted. Together, these results support a simple principle: allocate compute according to predicted incremental return, not difficulty alone.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.