SudokuRAG: Efficient RAG Prefilling via Coordinated Intra- and Inter-Document Attention Sparsity
Abstract
Retrieval-augmented generation (RAG) can incur substantial prefill overhead when multiple retrieved documents are incorporated into the model input. Existing approaches accelerate RAG prefilling through parallel encoding, but often require additional training or incur accuracy degradation. We observe that intra-doc attention patterns can be profiled offline and reused across prompts, whereas inter-doc selection must adapt to the current context. Building on this insight, we introduce SudokuRAG, a training-free design that coordinates offline intra-doc masks and online inter-doc pruning within the tiled attention loop, using current prefix and intra-doc softmax statistics as guidance. Empirical results show that SudokuRAG achieves prefill-attention speedups over dense attention at approximately 50% sparsity while largely preserving task accuracy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.