acceptodds
Under review as a conference paper at ICLR 2027

Iceberg Attention: Sublinear Long-Context Decoding with Mixture-of-Gaussians

Abstract

Autoregressive decoding over long contexts repeatedly scans the entire KV cache, though only a small fraction of cached tokens strongly influence each query. We introduce Iceberg Attention, a class of attention mechanisms that exploits this asymmetry by attending exactly to a small, query-dependent subset of the KV cache while analytically approximating the contribution of remaining tokens. Unlike sparse or retrieval-based attention, Iceberg Attention does not discard the residual cache; instead, it represents its aggregate contribution using a compact approximation. This formulation decouples the cost of exact attention from the context length and enables sublinear KV-cache access. We present Mixture-of-Gaussians (MoG) Attention, the first Iceberg Attention mechanism. % After prefill, MoG hierarchically partitions the keys of each attention head into clusters and subclusters per cluster, representing each group by a Gaussian distribution. During decoding, the query is scored against these compact representations to identify the most relevant clusters and subclusters, which are attended exactly. The remaining cache is approximated from the means and variances of its Gaussian representations. MoG reduces per-token attention from to without retraining or fine-tuning the model. We implement MoG as a drop-in replacement for attention in vLLM with a custom CUDA implementation that substantially reduces KV-cache data movement. At a context length of 1M tokens, MoG decoding is 27 faster than vLLM with FlashAttention and 2 faster than RetroInfer, a state-of-the-art sparse attention system, on an Nvidia H100 GPU. Across eleven long-context benchmarks using Qwen3-4B, Olmo3-7B, and Qwen3-14B at 8K–128K tokens, MoG matches exact attention while reading less than 5.3% of the KV cache. Ablations show the residual is what makes the method safe, cutting error of attending the top clusters alone from to . These results demonstrate that long-context decoding can avoid reading most of the KV cache without modifying or retraining the underlying model.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.