acceptodds
Under review as a conference paper at ICLR 2027

Rewards for budget-aware multi-turn agents that dynamically adapt to arbitrary test-time constraints

Abstract

Large language model agents may be given an inference budget and asked to maximise success within it. Budget-aware training has largely focused on controlling the cost of a single model generation at a single budget setting. However, multi-turn agents face a different problem, as a shared budget must be allocated across multiple environment interactions. This allocation is often managed by an external harness, adding inference-time complexity while leaving open whether the agent itself can learn to manage its own resources. We ask whether agents can be trained to reason about user-specified budgets and adapt when and how much they spend to maximise task success within them.We adapt existing length penalties to the multi-turn setting, propose new alternatives, and systematically study which objectives best teach agents to manage their resources effectively. We train Qwen3.5 models at 4B and 9B on AppWorld with budgets over tokens, API calls, and environment interactions, and evaluate generalisation to unseen budgets, to distribution shift, and to a second agentic benchmark. Training substantially improves both within-budget success and budget compliance over untrained models, and a single policy generalises this behaviour zero-shot to budgets, shifted distributions, and environments it never trained on, adjusting how much it spends to both the stated budget and the difficulty of the task. Simple gating, which withholds reward past the budget without pricing overspend magnitude, is highly competitive, suggesting that precise reward-shaping is not required for controlling overspend severity. Budgeting across multi-turn interactions can therefore be learned by the policy itself, without the external harnesses, predictors, or routers that prior approaches rely on.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.