acceptodds
Under review as a conference paper at ICLR 2027

Mechanics of Long-Context Hybrid Models 1.1: From Hybrid Attention to Hybrid Position

Abstract

The architectural design of Large Language Models (LLMs) is shifting from traditional full-attention-only models to hybrid models, which combine different attention modules to improve long-context efficiency and performance in length extrapolation and context extension. To explain why hybrid models work and how to design them better, we propose **Mechanics of Long-Context Hybrid Models**. As Part 1.1 of this series, we begin with hybrids of full attention and either sliding-window attention (SWA) or gated variants of linear attention (LA), represented by GLA and GDN. We first observe a **Seesaw Effect in Context Extension**: SWA hybrids perform better in length extrapolation, while LA hybrids perform better after context extension. We interpret this from the perspective of hybrid position. We find that SWA hybrids suffer from a **Short-Context Learning Trap**, **Short-Window Weariness**, and **Long-Window Laziness**, requiring extended windows to enhance performance in continual long-context pretraining. In contrast, we summarize the **Matthew Effect of Hybrid Position Extrapolation** and propose **Sliding-Window Linear Attention** for LA hybrids, achieving up to 16 training-free length extrapolation with 100% retrieval accuracy on NIAH-SK1.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.