acceptodds
Under review as a conference paper at ICLR 2027

CISA: Retrofitting Dense Attention Models with Chunk-Indexed Sparse Attention

Abstract

As context becomes grows, LLM inference becomes increasingly expensive because attention repeatedly reads a growing key-value (KV) history. We introduce Chunk-Indexed Sparse Attention (CISA), a method for retrofitting pretrained dense attention models without gradient updates. CISA selects fixed-size chunks of the KV history using compact summaries derived from the model's existing keys, avoiding token-level indexing and a separate full-length index cache. For indexing, CISA uses a single query per KV group and shares the selected chunks across the query in that group. During prefill, we reuse chunk selections across adjacent queries to reduce indexing cost and improve KV reuse and GPU utilization. We call this technique *QShare*. We evaluate CISA on the NVIDIA Nemotron 3 Nano and Ultra models, as well as on Inkling-Small. In one-layer BF16 benchmarks at batch size 32 using Ultra's per-device attention dimensions, CISA accelerates core attention by for prefill queries near the end of a 1M-token prompt and during decode. Including chunk scoring, selection, and all other measured GPU work, the corresponding speedups are and . On Ultra NVFP4, at , CISA keeps aggregate scores within 1.5 percentage points of dense attention on MRCR, LongBench v2, and GraphWalks. RULER sweeps on Nano and Ultra BF16 up to one million tokens show scores within 2.05 percentage points of dense attention with 8–16 fewer KV reads. When serving million-token prompts with 32 concurrent requests and 4K output tokens per request, CISA achieves the end-to-end throughput of dense attention using Nemotron 3 Ultra NVFP4 on four NVIDIA B300 GPUs. Selecting fewer chunks increases throughput to compared with dense attention; in separate GraphWalks evaluations, the smaller budget widens the aggregate score gap from 1.49 to 3.59 percentage points. These results expose a practical quality–speed tradeoff without requiring sparse training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.