acceptodds
Under review as a conference paper at ICLR 2027

MemToken: Scaling Query-Agnostic Context Compression across Tasks and Context Lengths

Abstract

Long contexts in real-world applications make the inference of Large Language Model (LLMs) expensive since attention prefill and key–value cache costs grow with sequence length, which motivates the research on context compression. Soft token compression aims to encode a long context into a short sequence of continuous vectors that can be directly inputed into the LLM and replace the original context. Existing query-agnostic soft compressors are typically trained on limited data distributions and relatively short contexts, leaving generalization to unseen data and longer sources unresolved. To address these limitations, we introduce MemToken, a query-agnostic soft token compressor that condenses groups of context representations into memory tokens consumed by a frozen LLM. To improve distributional generalization, we construct a broad-coverage training set that combines web-text reconstruction with synthetic supervised samples across diverse downstream tasks. To scale source length, we use a short-context reconstruction warmup, interleave short and long contexts, and fine-tune on longer task contexts. Across four unseen long-context benchmarks with different LLMs, MemToken scales the compressed context length to 128K and significantly outperforms the state-of-the-art methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.