MANTA: A Multi-level Aggregation Transformer
Abstract
Large language models are increasingly deployed on tasks that require access to large amounts of context, including document collections, code repositories, and long-running agent trajectories. Processing such extensive information with transformers is computationally and memory intensive: during autoregressive generation, the key-value (KV) cache grows linearly with sequence length, increasing accelerator memory demands and limiting throughput. The challenge is therefore to reduce the memory required to represent the full sequence history without compromising the information available for future predictions. We introduce MANTA, a Multi-level Aggregation Transformer that progressively aggregates distant context while retaining recent tokens at full resolution. MANTA maintains a window size at each aggregation level, resulting in logarithmic KV-cache growth while retaining information spanning the full sequence history. We evaluate MANTA by adapting the Llama 3.2 1B base architecture and compare it with a dense-attention baseline. Despite its substantially smaller KV cache, MANTA achieves competitive language-modeling performance and retains strong performance on long-context tasks. In a typical configuration, MANTA reduces the KV-cache size by 97.5% at a context length of 8K tokens. We further show that, with a natural modification, aggregation can replace tokenization, and that MANTA can be readily adapted to perform nearly lossless token compression. Overall, multi-level aggregation offers substantial architectural flexibility while directly addressing the memory capacity and bandwidth bottlenecks that increasingly constrain large-scale LLM inference.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.