QUOTA: Learning Downstream-Aware Context Allocation for Tool-Using Language Agents
Abstract
Tool-using language agents accumulate large volumes of tool outputs over long-horizon interactions. Existing compression methods primarily optimize immediate context length, but removing critical information can trigger task failures or redundant tool re-invocations, causing local token savings to increase end-to-end cost. We introduce QUOTA, a query-conditioned framework that learns how to allocate context budgets across individual tool outputs. At each compression event, QUOTA assigns a continuous retention quota to every output based on the query, execution state, and output characteristics. The policy is trained with REINFORCE using an objective that jointly considers task success, agent token consumption, and post-compression tool re-invocations, while keeping the agent and compressor frozen. To address delayed credit assignment, QUOTA learns an outcome model that compares each selected quota against a full-retention counterfactual, producing output-specific signals for policy optimization. Across held-out queries from a production data-analysis agent and three compressor backends, QUOTA reduces agent token consumption by 29%-46% relative to full retention, at success rates of 0.75-0.81 against 0.75 for full retention. It adds 15%-18% savings over a tuned tool-type-conditional policy. Interventional rollouts validate the learned outcome estimates, while cross-backbone and cross-compressor experiments demonstrate policy transfer. A separately trained policy reduces agent tokens by 18%-25% on five unseen LOCA-bench environments at success rates of 0.38-0.48 against 0.37 for full retention, showing that downstream-aware context allocation improves the efficiency of long-horizon language agents.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.