acceptodds
Under review as a conference paper at ICLR 2027

TaskAnchor: Evaluating Context Compaction in Long-Horizon Coding Agents

Abstract

Context compaction is essential for modern coding agents to handle practical long-horizon tasks, but current evaluations leave three gaps. First, evaluations with 32K/64K context budgets test compaction on relatively short inputs, leaving its effectiveness on histories spanning hundreds of thousands of tokens insufficiently evaluated. Second, evaluations on standalone texts overlook how coding agents select history for compression, retain recent messages, and recover omitted information through tools. Third, binary task rewards are too coarse to reveal information loss and behavioral changes after compaction. To bridge these gaps, we introduce TaskAnchor, a compaction-evaluation framework that compares information accessibility and subsequent behavior from shared execution states in native coding agents. Guided by TaskAnchor's evaluation, we propose StateCarry, a training-free compaction method combining context-aware tool-output compression with source-grounded summary revision to preserve evidence with its scope. Experiments across four scaffolds reveal substantial information loss under native compaction, while comparisons across 20 tasks on OpenCode show that information accessibility, behavior, and task outcomes can diverge. At approximately 200K-token checkpoints, StateCarry reduces context by 81.9% and improves the Strict score over native compaction by 14.9 percentage points, with few adverse events identified by automated evaluation in short continuations. In end-to-end runs, it shows no decrease in solve rate or fail-to-pass (F2P) test passes relative to native compaction.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.