acceptodds
Under review as a conference paper at ICLR 2027

IGAL: Importance-Guided Attention Linearization of Large Language Models

Abstract

The quadratic cost of self-attention makes long-context processing expensive for pretrained Transformers. Converting self-attention to more efficient alternatives, such as sliding-window attention (SWA) and linear attention (LA), reduces this cost but can degrade accuracy, particularly on tasks requiring long-range retrieval. Attention heads exhibit substantial functional heterogeneity, with different heads specializing in local interactions and long-range retrieval. Existing conversion methods typically approximate each pretrained attention head with a combination of LA and SWA, relying on head-specific mixing weights and/or branch parameters to adapt this hybrid form to different functional roles. In contrast, we leverage pretrained head specialization to directly determine the attention operator assigned to each head. Specifically, we propose Importance-Guided Attention Linearization (IGAL), which repurposes existing head-importance scores—previously used to allocate larger KV-cache budgets—to guide heterogeneous attention conversion. IGAL converts a small subset of high-scoring heads to LA for long-range modeling, while converting the remaining heads to SWA for local interactions. We further find that the optimal proportion of LA heads differs between short- and long-context tasks. To accommodate both regimes within a single checkpoint, we introduce IGAL-R, which sets the LA budget using the prompt length and learns which heads receive it with a lightweight prompt-level router, fine-tuned on a mixture of general-language and retrieval data. Under task-specific fine-tuning on Llama 3.1 and Qwen3, IGAL improves general-language and long-context retrieval performance over the existing state-of-the-art linearization methods. On Llama-3.1-8B, assigning only of GQA groups to LA achieves needle-in-a-haystack retrieval accuracy at 32K after 1K-length data fine-tuning, with a smaller decoding cache than full attention. Under matched mixed-data fine-tuning, IGAL-R raises 16–32K retrieval accuracy from to over static IGAL, with comparable general-language performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.