How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks
Abstract
Wide adoption of AI agents in complex human workflows drives rapid growth of LLM token consumption. When agents are deployed on tasks that can require a large amount of tokens, three questions naturally arise: (1) How do AI agents spend the tokens? (2) What models are more token efficient? and (3) Can LLMs anticipate the token usage before task execution? In this paper, we present a systematic study of token consumption in AI agents, analyzing up to twelve frontier LLMs across two coding-agent harnesses and four agentic environments. We find that: (1) agentic tasks are uniquely expensive, with input tokens rather than output tokens accounting for most of the consumption. (2) Token usage is highly variable: on the same task, the most expensive run typically consumes twice as many tokens as the cheapest, and higher usage does not guarantee higher accuracy. (3) Human-labeled task difficulty only weakly aligns with actual resource expenditure. (4) Models differ systematically in token efficiency, with similarly accurate models consuming up to 3× different amounts, but the harness can change consumption several-fold and even reorder models. (5) Finally, agents show promise in anticipating their own cost: the strongest models' pre-execution predictions correlate with actual usage at up to 0.77, a potential signal for budget alerts, although they tend to underestimate input tokens. Our study reveals important insights regarding the economics of AI agents and can inspire new studies in this direction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.