acceptodds
Under review as a conference paper at ICLR 2027

DySCO: Dynamic Attention-Scaling Decoding for Long-Context Language Models

Abstract

Understanding and reasoning over long contexts is a crucial capability for language models (LMs). Although recent models support increasingly long context windows, their accuracy often deteriorates as input length grows: models struggle to keep attention focused on the most relevant context throughout decoding. We propose DySCO, a training-free decoding algorithm for long-context reasoning. At each decoding step, DySCO uses retrieval heads–a small subset of attention heads specialized for long-context retrieval—to identify task-relevant tokens, and explicitly up-weights attention to them across all heads. DySCO applies directly to off-the-shelf LMs and is compatible with FlashAttention, as its intervention reduces to a per-token logit bias inside the attention kernel. Across three model families, spanning instruction-tuned and reasoning models, DySCO consistently improves performance on challenging long-context reasoning benchmarks, with relative gains of 22% and 21% for Qwen3-8B on MRCR and LongBenchV2 at 128K context length. DySCO incurs about 1.8x decoding latency but makes more effective use of test-time compute, outperforming self-consistency baselines with up to 8 parallel samples. Further analysis shows that both dynamic rescaling and retrieval-head guided selection are important to its effectiveness, and provides insights into decoding-time attention behavior.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.