Forgetting without Forgetting: Query-Dependent Decays for Effective Long-Context Modeling
Abstract
Position encodings are an integral part of transformers, heavily influencing language modeling performance as well as capabilities like recall and extrapolation to long contexts. Yet no existing encoding excels at all three aspects. The ubiquitous encoding RoPE performs well within training lengths, but degrades beyond. Decay-based encodings (like ALiBi or FoX) can extrapolate better, as they favor recent tokens more strongly, but this hurts long-range recall. Removing position encodings entirely (NoPE) can improve recall, but hurts language modeling. To resolve this, we introduce Query-Dependent Decay (QDecay): a lightweight component that can be added on top of any decay-based encoding, enabling each query to independently control the intensity of the underlying decay. A query can choose to either favor recent tokens (like ALiBi), or "turn off" the decay completely when distant information matters (like NoPE). Critically, this choice is made independently for each query in the sequence – favoring recent tokens for one query does not affect the decay for the next. Attaching QDecay on top of the simple ALiBi decay, the resulting position encoding achieves strong performance on all three desiderata. On standard language modeling benchmarks, QDecay matches or outperforms other position encodings, including the popular RoPE. Next, QDecay achieves the best recall on RULER at 32K+ token context, improving on other encodings by 18+ accuracy points, and remains competitive at shorter lengths. Third, QDecay extrapolates well: on Books3, it matches or outperforms other encodings on validation perplexity at all context lengths, maintaining and even improving performance through 64K tokens (16× training length), whereas all other encodings degrade. Our interpretability ablations identify QDecay's ability to completely "turn off" the underlying decay as a key driver of recall. QDecay adds negligible (< 0.1%) overhead in parameters and is compatible with fast FlashAttention-based algorithms.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.