NCSA: Native Chunk Summary Attention
Abstract
Long-context language modeling faces two bottlenecks: attention computation grows rapidly with context length, and the key–value (KV) cache grows with every token during generation. Recent KV-compression methods address these bottlenecks by representing several tokens with a compact summary. Existing designs mainly follow two approaches. Dedicated-compressor designs generate compressed KV entries through an additional module and computation path. Summary-token designs instead insert a learned summary for each chunk and process it as an ordinary sequence position. Consequently, the extra positions pass through every layer’s feed-forward network, and a summary produced in one layer is not available to text attention until the next. In this work, we introduce NCSA (Native Chunk Summary Attention), a nearly parameter-free architecture that generates and uses chunk summaries within the same attention layer. NCSA keeps summaries outside the token sequence in a summary residual stream, thereby eliminating the one-layer delay and making summary-side feed-forward computation optional while preserving the compressed KV cache. At the 4K context length used for pretraining, NCSA reduces estimated Transformer-block FLOPs by 14.2% relative to the summary-token baseline. Across matched pretraining runs at 0.3B, 0.6B, and 4B parameters on 50B tokens, NCSA consistently improves long-context retrieval over the compressed-attention baselines while remaining competitive in language modeling and commonsense reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.