acceptodds
Under review as a conference paper at ICLR 2027

What Compression Preserves: Instruction Survival in Long-Horizon LLM Agents

Abstract

When a long-running LLM agent fills its context window, production frameworks compress its history and continue. This compaction becomes the agent's effective memory of its own session, but it is unclear which standing instructions survive the rewrite. Across 1,175 evaluable trajectories on -bench retail and AppWorld, we vary three compression trigger thresholds, two compression methods (truncation and LLM summarization), and three reminder frequencies against a no-compression baseline, mechanically scoring every planted instruction. We find selective forgetting based on whether later trajectory content re-cues a requirement: on -bench, two task-external output requirements stated in the first user message drop from 82% and 87% without compression to 22% and 32% across compressed conditions, whereas rules coupled to the ongoing trajectory stay at 89–100% when pooled over all conditions. Among compressed trajectories that complete the task, 63% and 55% still violate these two requirements, so task success alone does not measure memory faithfulness. The pattern persists under changes to compaction settings and in a 40-task pilot with a second agent model family. The most aggressive compression also reduces task success by up to 24.7 (-bench) and 33.6 (AppWorld) percentage points and may increase input-token use: under uncached full-schema billing, compressed -bench conditions use 13–42% more input tokens than baseline; caching the schema prefix largely removes truncation's excess but not summarization's. Scheduling a single mid-conversation reminder improves compressed -bench adherence by 5.6 percentage points, an effect that does not replicate on AppWorld. We release the evaluation harness, constraint catalogue, and trajectories.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.