What Does Context Compression Cost an Agent? Interaction Costs Unrevealed by Task-Completion Metrics
Abstract
Task completion is the standard metric for evaluating context compression, yet it is an incomplete measure: compression can substantially increase an agent’s interaction cost — the reacquisition of state it dropped — without any detected change in completion. We frame completion-only evaluation as a lossy projection of the full evaluation outcome and introduce a controlled runtime protocol that measures this reacquisition cost in a bounded-horizon agent. Across three models and two task regimes, retrieval tool calls increase in every one of six model–regime comparisons, five of them after Holm correction, while execution calls change little. At the pre-specified 5× point no completion change is detected in any cell, whereas the retrieval signal already appears: on GPT-5.6 Sol completion is unchanged ([−5, +6] pp; MDE 8.4 pp) while retrieval rises by 24.2 calls. At more aggressive compression, completion degrades as reacquisition consumes the horizon. Oracle restoration and retention interventions probe the mechanism: restoring the queryable dropped state removes most of the added tool cost, while replacing D content with fabricated irrelevant state raises retrieval without a detected completion change. Random retention does not differ detectably from an offline hindsight oracle, suggesting state validity and presence matter more than fine-grained selection in this regime. The effect is not universal: fact-preserving extractive and abstractive operators avoid the surge, while the dropping operator produces it, and the surge reappears in τ-bench; the ALFWorld probe dropped no context and is uninformative about compression. Context compression can therefore incur hidden interaction cost when execution-relevant state must be reacquired, and completion alone can miss that cost.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.