acceptodds
Under review as a conference paper at ICLR 2027

DECAF: Cost-based Dynamic Agentic Planning under Budget Constraints

Abstract

Test-time scaling and structured planning let LLM agents tackle long-horizon tool use a single forward pass cannot, but every reasoning step, candidate plan, and re-plan costs tokens and dollars. Under a bounded budget an agent can spend so much on preparatory thinking that it never reaches the answer: the task is solvable, yet the run ends with no response because the budget is gone. This failure, which we call budget exhaustion, is invisible to how we price agents, which counts tool calls and answer tokens but treats the reasoning and planning that drive them as free. Avoiding it is hard because spending well is three coupled decisions (how much to plan, when to re-plan, and which path to commit to), each depending on a task's difficulty, learned only by acting, and on the budget left. Existing structured planners fix these and are cost-blind; the closest plan-graph planners draft a fixed number of plans, re-plan on every mismatch, and carry no notion of price. We propose DECAF (Dynamic, Efficient Cost-control for Agentic planning under a Finite budget), a plan-graph method that instead makes all three decisions budget-aware: it scales breadth to observed difficulty, gates each re-plan by a priced divergence test, and selects the executed path with a priced graph-ILP that prefers the cheapest goal-reaching route. On five benchmarks DECAF gives the best accuracy-cost trade-off among structured planners: no baseline is both cheaper and more accurate, improving accuracy by points on average over the strongest structured-planner baseline at its per-task cost; and on AppWorld it reaches at half the base budget the accuracy the strongest baselines need three times the budget to match.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.