CliffCompaction: Cost-Efficient Compaction for Long-Horizon Coding Agents
Abstract
Agents often work on complex problems that require millions of tokens of context, necessitating compaction across sessions with limited context windows. We develop CliffCompaction, an autocompaction technique that reduces cost by up to 50% under a bounded context while maintaining or improving performance on tasks such as Terminal-Bench. Its per-rollout savings can be reinvested in test-time scaling, improving Terminal-Bench by over 10 points for under the cost of two full-context runs, letting Kimi K2.6 match Opus 4.7 and exceed Opus 4.6 and GPT-5.3 Codex at lower cost. On KernelBench, CliffCompaction achieves state-of-the-art CUDA kernel speedups ( at 200 steps, at 400) across sessions exceeding a million tokens, surpassing trained agents and specialized search methods. The key to CliffCompaction is preserving context fidelity: (1) it only truncates or drops content, never rewrites it, and (2) never re-compacts a compaction, preventing context drift. We open-source a scaffold-agnostic API-proxy implementation usable with Claude Code, Codex, and other harnesses.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.