ReConBench: Evaluating Context Compression through Executable Agent Resumption
Abstract
During a long task, an agent accumulates a history of decisions, tool calls, and observations. Compressing this history reduces the context passed to a successor, but may also remove information needed to finish the task. We introduce ReConBench, a benchmark that evaluates context compression through task completion after an interruption. Starting from a verified successful execution, we select a checkpoint, restore its workspace, compress the preceding interaction history into a memory, and ask a fresh agent to complete the task. A task-specific verifier evaluates the final workspace. The benchmark pool contains 2,694 successful trajectories organized into four workflow families. Our main study evaluates 32,400 continuation runs across interruption stages, compressor scales, memory methods, and budgets within three source/continuation-agent strata. Among verifier-valid runs, 57.3% complete the task, with a mean normalized score of 0.834. Success rises from 49.4% at an early checkpoint to 67.8% at a late checkpoint, and from 54.2% with a 2B compressor to 59.6% with a 9B compressor. Memory methods differ by only a few percentage points, and larger allowances produce neither proportional increases in memory length nor consistently higher success. These findings show why compression should be evaluated through both the task it preserves and the work required after handoff.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.