acceptodds
Under review as a conference paper at ICLR 2027

QKSieve: Query-Key Balanced Mixed-Bit Retrieval for Efficient Long-Context Attention

Abstract

Autoregressive decoding repeatedly scans the complete key-value (KV) history. Query-aware sparse attention helps only when selection is cheaper than dense QK scoring and faithful to joint Query-Key geometry. We present QKSieve, a training-free, GPU-resident exact-KV retrieval method. Request-local Query and Key second moments define biorthogonal coordinates that preserve every unquantized dot product and order dimensions by joint score energy. A per-layer, per-head allocator assigns 0/1/2/4/8 bits to eight 16-D bands under a 240-bit token/head budget by minimizing a separable logit-MSE surrogate. The packed index targets B(N) = minN, 1280, max(256, ceil(0.06N)) positions, after which the original 16-bit K/V receive exact sparse attention. No learned router, exact-QK reranking, task rule, or dense fallback is used. QKSieve-Robust adds a current-Query rank-16 INT4 estimate of the omitted softmax partition and Value numerator, bringing total auxiliary storage to 7.47%. On complete LongBench, Robust retains 99.98%, 99.73%, and 100.00% of Full macro score on Llama-3.1-8B, Qwen3-4B, and Mistral-7B while targeting 6.85-7.43% of history. Across 13 RULER tasks from 4K to 128K, it retains 100.33% on average with 4.97% measured active attention. In a controlled native-MHA RTX 3090 implementation, the complete Robust attention path is 2.09-4.12x faster at 32K-128K, and real steady decode is 1.32-3.98x faster. We separately report index construction and persistent-cache break-even rather than hiding fixed cost in steady-state timing. Analysis connects representation error to ranking margins, omitted probability mass, and attention-output error.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.