acceptodds
Under review as a conference paper at ICLR 2027

R-Notes: Sliding-Window LLMs Can Learn to Remember Beyond Their Attention Window

Abstract

Real-world LLM deployments increasingly require processing inputs of effectively unbounded length, calling for models that can handle arbitrarily long inputs with bounded memory. However, existing constant-memory architectures, such as linear attention and pure sliding window attention (SWA), lose information beyond their fixed state or window. To address this, we observe that an SWA model can preserve information much like a reader taking notes: brief self-written notes stay in the recent window and carry key content forward, even as the original tokens scroll out. Trained to do this consistently, pure SWA models already suffice for lengthy inputs. Building on this observation, we propose R-Notes, which feeds fixed-size chunks of the input as a multi-turn conversation under a single SWA session and trains the model to emit brief notes that propagate key information across chunks within the window. Built on Qwen3-8B, R-Notes matches or exceeds the 1.75x larger Qwen3-14B full-attention baseline on LongBench cross-task QA (73.4% vs 71.3%) and on the longest settings of RULER-HQA at 112K and RULER-SQuAD at 128K (87.50% vs 78.12% on both), while keeping the KV cache bounded by the window size and total compute linear in the number of processed chunks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.