acceptodds
Under review as a conference paper at ICLR 2027

GLIDE: Guided Layerwise Hybrid Attention for Efficient LLM Inference

Abstract

As Large Language Models scale to increasingly long contexts, the memory I/O and computational overhead of the Key-Value (KV) cache during decoding emerges as the primary throughput bottleneck. To address this, we propose GLIDE, a Guided Layerwise Hybrid Attention that strategically integrates sliding-window softmax attention with linear recurrent aggregation. GLIDE is motivated by layer-wise heterogeneity: early layers exhibit high sensitivity to softmax removal, while deeper layers demonstrate redundancy and tolerate aggressive replacement by linear alternatives. Leveraging this insight, GLIDE introduces a layer-wise adaptive mechanism wherein each layer is balanced with fully variable configurations between efficient linear recurrence and full-accuracy softmax. Employing front-loading softmax retention strategies, GLIDE non-uniformly compresses the softmax allocation in the model, reducing aggregate KV cache I/O while retaining model accuracy in sensitive layers. We demonstrate GLIDE’s framework adaptability and superior performance-efficiency tradeoffs with multiple linear and hybrid attention models, reducing end-to-end latency for long-context generation without compromising quality.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.