Spend Less, Reason Better: Budget-Annealed Tree Search for LLM Agents that Spend Only on New Evidence
Abstract
Test-time scaling lets tool-augmented LLM agents trade compute for accuracy, but much of that compute is wasted: agents re-issue tool calls they have already made, follow dead-end trajectories, and exhaust their tool and token budgets before reaching an answer. Addressing this waste requires jointly allocating the remaining budget across intermediate reasoning states and reusing evidence across branches, rather than treating budget control and evidence handling as separate decisions. We propose the Budget-Aware Value Tree (BAVT), a training-free inference framework that organizes reasoning and tool use as a tree scored by a step-level critic from the same backbone, made budget-aware by two mechanisms. Budget-annealed selection sharpens the sampling distribution over candidate nodes by the inverse of the remaining budget ratio, so the agent explores broadly while resources are ample and concentrates on the most promising branch as they run out. Evidence reuse charges the budget only for new evidence: same-parent exact repeats are merged before execution, while other eligible store hits reuse cached evidence and are re-scored in their new branch context, and every saved call is returned to the budget controller. For a capped selection variant, we prove a conditional finite-budget guarantee for generating a terminal answer. Across four multi-hop QA benchmarks, two model families and three budget tiers, BAVT outperforms parallel sampling, ReAct, LATS and BATS at every budget, using fewer paid tool calls than parallel sampling. With a five-call budget, it surpasses baselines given a larger budget while using about one fifth as many paid tool calls as parallel sampling; instantiated for browser agents on WebArena-Lite, it leads all baselines at every action budget, and the ordering carries over to three closed-weight models on both tasks. Allocating a fixed budget adaptively is thus more effective for reliable agents than brute-force test-time scaling.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.