Learning Optimal Agents Across Any Budget via Bilevel Optimization
Abstract
Autonomous agents often achieve stronger task performance by consuming more execution resources, leading to increased deployment cost. To improve efficiency, existing budget-aware approaches train agents under particular budget conditions or expose them to multiple predefined budgets.In practical deployment, however, the available budget can vary substantially across scenarios, motivating agents that can perform well across the entire budget range. In this work, we study how to train an agent that perform reliably across varying budget constraints. We transform the objective of maximizing any-budget capability into minimizing the expected minimum sufficient budget required by individual tasks, which naturally yields a coupled bilevel optimization problem over task-specific budget adaptation and policy improvement. Based on this formulation, we propose BABLE (Bilevel Any-Budget Learning for agEnts), which solves the coupled problem through alternating optimization. BABLE combines a Task-Adaptive Budget Controller (TABC), which adapts task-specific budgets from recent successful executions, with Likelihood-Guided Conditional Advantage Optimization (LCAO), which strengthens policy learning from sparse successes under tight budget conditions. Experiments on four agentic benchmarks under both token-level and tool-call level budgets show that BABLE consistently improves performance across the evaluated budget range compared with standard reinforcement learning and budget-aware baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.