MANTA: Information-Adaptive Context Compression via Matryoshka Tokens and Tolerance-Conditioned Allocation
Abstract
Context compression enables large language models to process long inputs more efficiently by encoding the original context into compact, reusable representations. However, most existing methods fix a single compression ratio at training time, while a few support several predefined ratios. Either way, every context of a given length receives the same allocation, regardless of its information content, resulting in inefficient resource utilization. To address this, we introduce information-adaptive context compression: the allocation grows for information-dense contexts and shrinks for redundant ones. We formalize this through a criterion we call the minimal sufficient allocation: the smallest allocation whose excess loss, measured against the full allocation, remains below a global tolerance. We realize this criterion in MANTA, which exposes the compression ratio as a continuous inference-time condition, built from two components: (1) Matryoshka tokens, a compact nested representation of each chunk with valid arbitrary-length prefixes; (2) a lightweight Elastic Allocator, conditioned on the tolerance, that decides how many Matryoshka tokens each chunk retains from the excess loss its content would incur. MANTA consistently outperforms existing context compression methods across different compression ratios, improving average token-F1 by 9.3% relative in domain and 5.4% out of domain over the strongest compression-matched baseline. Further analyses attribute the improvements to information-adaptive compression, where the allocator flexibly distributes Matryoshka tokens according to information density.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.