acceptodds
Under review as a conference paper at ICLR 2027

SCOPE: Stored Chunk-Level Offline Priors for Efficient Long-Context Inference

Abstract

Long-context inference for large language models is limited by the repeated online scoring used to select important key-value (KV) blocks. We introduce SCOPE (Stored Chunk-Level Offline Priors), an efficient chunk-level method that reuses offline content priors across queries within a deployment domain. A store of offline chunk scores is built in the scheme of exponential moving average (EMA). Then unseen chunks are predicted with a lightweight ridge regressor, and finally intra- and inter-chunk refinement is applied at inference time with constant cost. This design replaces the dense scoring path of Flash Sparse Attention (FSA) with an expected- hash-table lookup per chunk followed by lightweight refinement. Although SCOPE uses static routing instead of the original query-dependent selector, it is still compatible with FSA’s sparse-kernel interface. SCOPE adds no trainable parameters and requires only a small amount of auxiliary storage. Its store construction and lookup path can also run on CPUs alongside GPU computation, reducing inference latency without extending GPU pretraining. On Qwen-3.5-2B and Llama-3.2-1B, SCOPE improves median latency in 31 of 32 released-backbone cells, with average median and P99 reductions of 11.5% and 12.8%, respectively. On the larger Qwen2.5-7B model with a 32K context, SCOPE achieves a 15.3% average median-latency reduction. These results show that offline chunk priors provide an efficient alternative to repeated online scoring for long-context inference.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.