Where to Recompute: Learning Recoverability for Compute-Efficient LLM Reasoning
Abstract
Test-time computation strengthens large language model reasoning only when the additional work changes the answer that the system ultimately returns, so deciding where and how much to recompute is central to compute-efficient inference. Existing methods face two challenges in making this decision. First, they target candidate correctness, continuation progress, or search coverage, yet a frozen selector can reject a repair or let a higher-scoring error replace a correct answer, so better generation does not by itself yield a better returned answer. Second, question-level allocation does not directly estimate the post-selection value of retaining a particular prefix and assigning it a continuation budget. To address the first challenge, we propose Recoverability-Aware Value of Computation(RAVEC), which learns the post-selection repair and damage of each prefix intervention from randomized same-parent interventions with a frozen generator and selector, and we prove a selector-gating identity under which changing only the selector can reverse the optimal intervention. To address the second challenge, we develop finite-horizon control over reopening positions, suffix budgets, and stopping, prove that path information improves on the best coarser policy exactly when its decision value exceeds estimation regret, and sharpen within-parent action ordering through parent-centered fitting without additional rollouts. Across 14 open models and 12 benchmarks, RAVEC improves core macro accuracy over adapted persistent-pool search (PB-SMC) by 1.34 points on average at equal expected operations, and at matched accuracy it reduces computation by 20.8% on MATH-500 and 25.4% on LiveCodeBench.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.