Towards Budget-Aware Agents: Do Agents Know What They Will Spend?
Abstract
Foundation-model agents operate under growing resource constraints yet rarely know how much budget they will spend. We call this capability **budget awareness** and formalize it as **progressive interval estimation**: mid-execution, can the agent provide an interval that covers its remaining budget and predict when its continuation will not complete within budget? We score this with a rollout-replay protocol that re-queries the agent on every trajectory prefix, decomposing estimation into feasibility prediction, early failure detection, and interval estimation. We evaluate five frontier models across four environments, including internal token budgets (Sokoban, Search-R1, SWE-bench) and external multi-dimensional budgets (Warehouse), and train Qwen-7B estimators with SFT and RL. Task success is only a loose proxy for estimation quality: frontier models differ significantly as estimators, but not in order of success. Estimation fails in structured ways: agents underestimate remaining cost on search and coding, recognize failure late, and widen intervals with task scale rather than continuation uncertainty. At a matched loss of solved tasks, the models’ infeasibility signal saves more tokens than a hindsight-tuned spending rule in 9 of 15 model–environment cells. SFT makes binary feasibility learnable in-domain, whereas interval estimation stays task-specific.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.