acceptodds
Under review as a conference paper at ICLR 2027

Oracle-Guided Sparse Prefill: Separating Oracle, Indexer, and Runtime Gaps in Long-Context GQA

Abstract

Long-context prefill remains expensive because global attention layers, including those using grouped-query attention (GQA), must compute pairwise query–key scores across the entire history, incurring quadratic cost in sequence length even in hybrid models that combine global attention with local, linear, or recurrent components. We study when this computation can be sparsified without materially changing task performance. We use an attention-mass top-k oracle that selects a head-averaged token support from dense attention. This non-deployable reference separates sparse-budget feasibility from indexer error and runtime realization effects. We evaluate a head-collapsed indexer trained by KL distillation with the backbone frozen. On Qwen3.5-27B RULER samples from 32K to 128K, matched per-query and 32-query-shared oracle and indexer configurations remain within one observed macro-accuracy point of dense. At 128K, the 32-query-shared indexer scores 0.61 points below dense with a 1.46 single-request TTFT speedup. On sampled tail queries, the indexer recovers 94–98% of the attention mass retained by the oracle, yet task-level failures coexist with this high aggregate fidelity. Key-block experiments show that oracle-feasible support structures can still fail under learned selection. These observations motivate configuration-aware support selection and establish a diagnostic study of sparse-prefill quality–latency tradeoffs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.