Token-Constrained Mission Completion in Agentic AI: Capacity Curves and Budget-Aware Reasoning Allocation
Abstract
An agentic system is given a task and an allowance. Benchmarks score the first and ignore the second: a run that would have been correct with twice the tokens scores exactly like one that was affordable. We make the allowance part of the state. Writing an architecture as a budget-augmented stochastic Petri net with state , the workflow marking and the remaining budget, makes the constraint structural: a transition the budget cannot pay for is disabled, so its call is never attempted, and exhaustion joins success and error as a third absorbing outcome. The capacity curve , the least budget attaining reliability , becomes computable rather than sampled. Four results follow. Accuracy does not order missions: a strictly more accurate reasoner is strictly worse over an interval of budgets exactly as wide as its price premium, so the inversion is generic, not a counterexample. Every fixed configuration has a reliability ceiling no budget crosses. Two capacity curves must cross whenever the higher-scoring configuration starts paying off later, so a score-ordered leaderboard is budget-invariant only under a condition no benchmark reports. And because spending is irreversible, the decision process is acyclic, so the optimal budget-aware policy is one backward pass. We then test the claims on 8,190 published agent trajectories. Reconstructing what each run spent turns every leaderboard entry into a capacity curve, and the ordering changes with the allowance: the top-scoring configuration has zero observed success below 23,126 tokens (exact paired test against a lower-ranked one, ) and 41 pairs cross. Selecting on capacity rather than score reaches the same reliability with 88% fewer tokens, and 45.7% of that benchmark's recorded failures turn out to be runs that produced no answer at all.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.