Characterizing and Optimizing Context Compaction in Production Coding Agents
Abstract
Long-running coding agents rely on context compaction to overcome finite LLM context windows. Existing harnesses typically perform compaction synchronously, placing expensive summarization on the critical path and disrupting prefix-cache reuse. We deep-dive into compaction behavior and costs by analyzing 13.5M GitHub Copilot agentic coding sessions from 3.2M users, comprising 760M LLM calls. Although only 7.8% of sessions trigger compaction, they account for 44.2% of total token consumption. Compaction also causes substantial foreground delay and sharply reduces prefix-cache reuse. To enable systematic evaluation, we first introduce COMPACTIONBENCH, a benchmark that models software evolution as persistent, multi-turn coding sessions. It enables controlled evaluation of repeated compaction while preserving executable task-level evaluation and reflecting context-growth patterns observed in production. We then present APEC, an asynchronous progressive compaction mechanism that prepares fine-grained, indexed summaries in the background and appends them to a stable prefix, moving compaction work off the critical path while preserving prefix reuse. APEC is harness-agnostic and treats summarization as a black box. Evaluated across frontier models, APEC improves task quality while reducing total cost by up to 25.4% and agent execution time by 9.6%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.