acceptodds
Under review as a conference paper at ICLR 2027

NCSA: Native Chunk Summary Attention

Abstract

Long-context language modeling faces two bottlenecks: attention computation grows rapidly with context length, and the key–value (KV) cache grows with every token during generation. Recent KV-compression methods address these bottlenecks by representing several tokens with a compact summary. Existing designs mainly follow two approaches. Dedicated-compressor designs generate compressed KV entries through an additional module and computation path. Summary-token designs instead insert a learned summary for each chunk and process it as an ordinary sequence position. Consequently, the extra positions pass through every layer’s feed-forward network, and a summary produced in one layer is not available to text attention until the next. In this work, we introduce NCSA (Native Chunk Summary Attention), a nearly parameter-free architecture that generates and uses chunk summaries within the same attention layer. NCSA keeps summaries outside the token sequence in a summary residual stream, thereby eliminating the one-layer delay and making summary-side feed-forward computation optional while preserving the compressed KV cache. At the 4K context length used for pretraining, NCSA reduces estimated Transformer-block FLOPs by 14.2% relative to the summary-token baseline. Across matched pretraining runs at 0.3B, 0.6B, and 4B parameters on 50B tokens, NCSA consistently improves long-context retrieval over the compressed-attention baselines while remaining competitive in language modeling and commonsense reasoning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.