acceptodds
Under review as a conference paper at ICLR 2027

From Rigid to Dynamic: Entropy-Guided Adaptive Inference for Accelerating Long-Context LLMs

Abstract

Existing sparse attention and KV cache compression methods for long-context LLM inference typically apply fixed sparsity patterns or uniform budgets across all attention heads, overlooking the substantial variation in attention behavior among different heads and contexts. We observe that attention heads exhibit two distinct entropy patterns: Rigid Heads, whose attention entropy remains near zero across different segment of inputs, and Dynamic Heads, whose entropy fluctuates significantly across the input context. Crucially, the distribution of these two types is context-dependent and cannot be predetermined offline. Based on this observation, we propose EntropyInfer, a training-free framework that leverages attention entropy to adaptively allocate computation budgets at the granularity of individual heads and segments during prefilling. For the decoding stage, we further introduce a latent KV cache compression scheme that utilizes generated output tokens, rather than prefill tokens alone, to identify and retain the most critical cache entries. Experiments on LongBench and InfiniteBench with Llama-3.1-8B-Instruct and Qwen-2.5-7B-Instruct show that EntropyInfer consistently outperforms existing methods including SnapKV, AdaKV and CritiPrefill, achieving up to 2.39 end-to-end speedup in scenarios where context length grows beyond 100k tokens, with minimal average generation quality degradation compared to full attention.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.