PickPack: Scaling Context Extension to Millions of Tokens with Sparse Attention
Abstract
Due to a growing range of applications, long-context processing and long-form generation with AI are becoming increasingly common, particularly with the transition to agentic workloads. Agents are long-lived, continuously produce and consume streams of tokens, and base their subsequent responses on their entire interaction history; multi-agent systems can further amplify this token volume. Models capable of processing long contexts are therefore increasingly necessary, but training with larger context windows is expensive, so most models are trained with windows of 32K to 256K tokens. Application demands can extend to millions of tokens, which creates a need for training-free context extension. Training-free context extension presents two key challenges: (1) out-of-distribution context embeddings, where unseen context lengths can shift attention distributions away from those seen during training, and (2) computational complexity, since attention's quadratic cost makes processing millions of tokens expensive. We propose , which addresses both challenges with two composable operators in a single FlashAttention pass: maps every query–key pair to a relative position inside the trained window, and uses the attention scores to dynamically decide which key blocks to skip. At 50% sparsity, improves accuracy over dense attention under the same position map on seven of eight Qwen3 model sizes and context lengths, while being up to 1.2 faster. Against the recipes that released Qwen2.5, Qwen3.5, and Qwen3.8 models ship with, it is 3.4–11 points more accurate and up to 1.2 faster from 128K to one million tokens, and 11.6 points more accurate at four million tokens, 16 the trained window of Qwen3.5-9B.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.