acceptodds
Under review as a conference paper at ICLR 2027

Prompt-Level Adaptive Memory Compression for Detecting Unreliable Memory

Abstract

LLM agents rely on long contexts held in working memory, from requirements retained by coding agents to evidence gathered by research agents. Compressing this memory lowers reading cost, but lost or altered details can make it unreliable for a given prompt, letting errors propagate through later actions. Existing compression methods treat compressed memory as a fixed input and offer little way to tell when it should not be trusted. We formulate prompt-level adaptive compression, in which an agent uses compressed memory while identifying, prompt by prompt, when it is unreliable, and propose Gated Context Memory (GCM). GCM stores context as continuous latent states produced by a writer and consumed by a reader that answers the prompt: a simple writer retains prompt-relevant information without converting it into text, and the confidence of a reader trained to use this memory can indicate when the memory falls short of what the prompt requires. The writer and reader share one frozen LLM with a small set of trained parameters. The trained reader supplies this signal at its first decoding step, without generating a complete answer, consulting an external verifier, or training a separate failure predictor, and the signal can drive any correction strategy. Across tool use, multi-hop reasoning, and long-document question answering, GCM outperforms existing compression baselines, and with raw-context fallback it matches or exceeds raw-context performance for most readers. Transfer experiments across tasks, readers, and compressors show where the signal generalizes and where it breaks down.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.