acceptodds
Under review as a conference paper at ICLR 2027

SudokuRAG: Efficient RAG Prefilling via Coordinated Intra- and Inter-Document Attention Sparsity

Abstract

Retrieval-augmented generation (RAG) can incur substantial prefill overhead when multiple retrieved documents are incorporated into the model input. Existing approaches accelerate RAG prefilling through parallel encoding, but often require additional training or incur accuracy degradation. We observe that intra-doc attention patterns can be profiled offline and reused across prompts, whereas inter-doc selection must adapt to the current context. Building on this insight, we introduce SudokuRAG, a training-free design that coordinates offline intra-doc masks and online inter-doc pruning within the tiled attention loop, using current prefix and intra-doc softmax statistics as guidance. Empirical results show that SudokuRAG achieves prefill-attention speedups over dense attention at approximately 50% sparsity while largely preserving task accuracy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.