Preview-Guided Memory: Learning to Reread for Long-Context Reasoning
Abstract
LLM agents accumulate long contexts as they read documents, use tools, and carry out multi-step tasks. Memory-based methods handle these contexts by reading text in chunks and maintaining a compact memory. However, single-pass readers must decide what to retain before seeing later chunks, while still processing every chunk regardless of its relevance. We propose Preview-Guided Memory (PGM), which uses a preliminary preview to guide both reading and memory updates. In the first pass, PGM generates a short, question-independent preview of each chunk and groups neighboring previews into pages. A learned selector uses these pages and the question to choose which chunks to read in detail in the second pass. The reader then processes the selected chunks in order, using the corresponding preview page when updating its memory. These previews provide clues about upcoming content, while selection reduces the number of memory updates. We keep the reader fixed and train the selector with supporting-fact supervision followed by reinforcement learning, rewarding answer quality and evidence coverage, with a bonus for reduced reading when the answer is correct. PGM-Selective-7B achieved state-of-the-art on both HQA and longbench v2 with fewer tokens for reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.