acceptodds
Under review as a conference paper at ICLR 2027

What Survives the Long Horizon Run? Benchmarking Task State in Coding Agents

Abstract

Long-horizon agents may run for hours to weeks, requiring automated context compaction to fit constrained LLM context windows. However, such compaction risks corrupting or obscuring critical task state: objective-relevant facts and their provenance. To address this risk, we introduce EndureBench to evaluate language agents' preservation of critical task state across compaction cycles. It builds histories from recorded coding sessions toward a 400K-token context target and compares 15 compressors at fixed checkpoints under two schedules: fresh, which independently compacts each raw history prefix once, and rolling, which recursively updates the previous memory with the next interaction block. The results are consistent across compressors: 14 of 15 lose quality under rolling, and the failures are outright misses rather than partial answers—current-state correctness scores 0 on 61.0% of judgments and scope binding on 38.1%, against 11.3% for verification calibration. Overall quality hides this failure: compacted memory may omit objective-relevant facts and their provenance and should be verified before the agent resumes rather than treated as authoritative. EndureBench thereby isolates state preservation as a distinct object of measurement for long-horizon agents.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.