acceptodds
Under review as a conference paper at ICLR 2027

LMI: LATENT MEMORY INTERACTION FOR EFFICIENT AND FAITHFUL LONG-CONTEXT COMPRESSION IN LLMS

Abstract

Large language models (LLMs) have demonstrated strong long-context understanding, yet scaling context length remains challenging due to the quadratic complexity of self-attention and the linear growth of the key-value (KV) cache. Existing compression approaches—token pruning, pooling, and hierarchical summarization—often suffer from irreversible information loss and limited global interaction among compressed representations. We propose LMI (Latent Memory Interaction), a window-parallel framework that partitions long sequences into context windows and introduces learnable latent tokens that aggregate and exchange information across windows. Unlike methods that independently compress each segment, LMI dynamically refines latent representations through cross-window interaction, preserving local detail and global semantic dependencies while maintaining near-linear complexity and compatibility with existing LLM architectures. Extensive experiments on long-context benchmarks demonstrate that LMI achieves competitive performance against existing token-reduction and compression methods; on a 128K-token, 32-request serving workload, LMI reduces end-to-end inference time by 8.7 relative to full-context inference, with the one-time compression overhead more than offset by the reduction in generation cost.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.