acceptodds
Under review as a conference paper at ICLR 2027

RouteVLM-NEXT: Making Vision-Language Models Look Smarter and Further Via Hybrid Attention

Abstract

Vision-Language Models (VLMs) are severely constrained by the quadratic complexity of self-attention, making the processing of high-resolution images and long videos computationally prohibitive. Existing approaches that uniformly compress or prune visual tokens often discard the very fine-grained details necessary for complex reasoning. To address this fundamental trade-off between efficiency and fidelity, we introduce RouteVLM-NEXT, a novel VLM architecture that pioneers a router-guided hybrid attention mechanism. At its core, a learnable router dynamically directs each visual token to one of two pathways: a standard full-attention branch for capturing intricate details in critical regions, or a highly efficient gated linear attention branch for processing background context. This mechanism, governed by a budget-aware training objective, allows the model to autonomously manage its computational resources, allocating high-fidelity computation only where it is most needed. Through extensive experiments, we show that RouteVLM-NEXT not only achieves excellent performance on a wide range of benchmarks, particularly excelling on detail-oriented tasks like V* Bench, but also exhibits great scalability. It can successfully handle sequences with up to 64k or even more labels, which standard full attention models cannot do on the same GPU, providing a new direction for the design of next-generation VLMs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.