acceptodds
Under review as a conference paper at ICLR 2027

GateSight: Recurrent Memory Anticipates Attention Memory for Query-Agnostic KV Eviction in Hybrid LLMs

Abstract

Hybrid language models interleave Gated DeltaNet (GDN) layers, whose recurrent memory has a fixed size, with full-attention layers, whose key–value (KV) cache grows with the context and dominates serving memory. In this paper, we ask whether the recurrent memory already knows what the attention memory will need. We first show that which GDN heads of Qwen3.8-27B keep long memories is fixed by the weights rather than by the document. Next, we define the future utility of a KV entry by how much removing it changes the attention output over the model's own continuations, and find that the write gate, decay and write magnitude that GDN layers emit during prefill anticipate this utility before any query is seen. Because the relation is spread across many GDN heads and layers, and a linear map recovers its trend but less of its high-utility tail than a nonlinear one, we propose GateSight, which learns it with lightweight per-head multilayer perceptrons that read these signals and no attention score. Their predicted utility decides which entries each attention layer keeps, and how much each layer keeps is set by the model's native attention output gate, whose mean differs by more than an order of magnitude across layers. Finally, at serving time, GateSight scores the cache once after prefill and physically releases discarded KV entries. Experimental results on SCBench and RULER-4096 with Qwen3.8-27B and, with the same recipe, the sparse-attention model Qwen3.8-Flash-Next, up to the 262K-token limit of the models, show that with a tenth of the attention KV cache, GateSight is the most accurate of five query-agnostic methods on three of the four model–benchmark pairs and leads KVzip on all four, by 9.5 to 47.2 points. Notably, it scores a full-length context with negligible overhead relative to prefill.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.