acceptodds
Under review as a conference paper at ICLR 2027

HyLin: Modality-Aware Hybrid Attention for Efficient Long-Context Vision–Language Models

Abstract

While large vision-language models scale to ultra-long sequences, the quadratic complexity of softmax attention becomes a severe bottleneck for both multimodal understanding and video generation. Existing inter-layer hybrids retain expensive full-attention layers, uniform intra-layer hybrids ignore modality-specific structure, and sparse video methods often preserve the full KV cache. We argue that efficient attention must combine global context compression with fine-grained local retrieval. Our teacher-attention analysis reveals substantial heterogeneity across modalities, heads, and depth: visual rows are often diffuse and compressible, whereas text rows favor local retrieval, and heterogeneous cross-modal mass makes universal replacement unsafe. We therefore propose \model, a Hybrid Linear architecture centered on Hybrid Linear-coupled Window Attention (HLWA). In replaceable VLM layers, HLWA routes visual queries through a constant-state global Linear Attention stream and applies Sliding Window Attention (SWA) directly to the original text queries, keys, and values before restoring multimodal row order. Theory-guided calibration and downstream sensitivity identify high-risk layers, with 25% retained as full attention. A four-stage curriculum performs operator alignment, layer selection, vision-language distillation, and supervised fine-tuning. For video DiTs, we instantiate the same global-linear/local-exact principle with a distinct bidirectional operator. On the reported benchmarks, \model-7B closely tracks its full-attention teacher while delivering prefill and decode speedup together with memory reduction at 64K tokens; the video-DiT variant accelerates HunyuanVideo-13B by with competitive VBench quality.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.