Oracle-Guided Sparse Prefill: Separating Oracle, Indexer, and Runtime Gaps in Long-Context GQA
Abstract
Long-context prefill remains expensive because global attention layers, including those using grouped-query attention (GQA), must compute pairwise query–key scores across the entire history, incurring quadratic cost in sequence length even in hybrid models that combine global attention with local, linear, or recurrent components. We study when this computation can be sparsified without materially changing task performance. We use an attention-mass top-k oracle that selects a head-averaged token support from dense attention. This non-deployable reference separates sparse-budget feasibility from indexer error and runtime realization effects. We evaluate a head-collapsed indexer trained by KL distillation with the backbone frozen. On Qwen3.5-27B RULER samples from 32K to 128K, matched per-query and 32-query-shared oracle and indexer configurations remain within one observed macro-accuracy point of dense. At 128K, the 32-query-shared indexer scores 0.61 points below dense with a 1.46 single-request TTFT speedup. On sampled tail queries, the indexer recovers 94–98% of the attention mass retained by the oracle, yet task-level failures coexist with this high aggregate fidelity. Key-block experiments show that oracle-feasible support structures can still fail under learned selection. These observations motivate configuration-aware support selection and establish a diagnostic study of sparse-prefill quality–latency tradeoffs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.