acceptodds
Under review as a conference paper at ICLR 2027

The Sink Tax: Auditing Memory Tokens and Assigning Their Roles in Long-Context Language Models

Abstract

In-context memory (ICM) systems reduce the key-value cache of long contexts by splitting the context into chunks, compressing each chunk into a few memory tokens, and evicting the raw tokens. However, whether the model uses the information that each memory token stores has not been examined. In this paper, we introduce a slot audit that tests whether a memory token stores the information of its chunk by replacing its key-value states with those of another document. We apply the audit to designs that evict the raw tokens, from models trained under the mask of gist and beacon tokens to the official checkpoints of AutoCompressor and RMT. We find that the attention sink shifts to the memory tokens, and across these existing designs the memory token that receives the most attention consistently stores no information of its chunk that the model uses. Yet, removing this token still raises the loss under the mask that hides each chunk from later queries. Therefore, the removal test, which existing work uses to validate memory tokens, counts the sink as memory. We term the fraction of memory tokens that the sink takes the sink tax. Based on these findings, we introduce Assigned-Role Memory (ARM), which assigns one role to every cache position, whereas in prior designs the roles emerge during training. ARM assigns the sink of the memory slots to one sink slot with a value fixed at zero and a token-like key. This sink slot removes the sink from every memory slot at 124M, 340M, and 0.8B parameters. Once the roles are fixed, ARM keeps one set of memory slots for the entire context and updates the memory slots after every chunk, whereas prior designs accumulate one set per chunk. The cache size therefore remains constant as the context length grows. In long-context language modeling on PG19, ARM matches the long-range perplexity of per-chunk accumulation with 17 cache positions instead of 384 and recalls 79% of the facts placed in earlier chunks compared to 54%. With 17 cache positions, ARM lowers the long-range perplexity of a sliding window by 29%. ARM maintains its perplexity at 128k tokens, 32× the training length, while per-chunk accumulation degrades beyond the training length. At 128k tokens, ARM runs at the decoding throughput of a sliding window with a lower perplexity and at 39× that of per-chunk accumulation at the ratio of prior work.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.