AdaVSkip: Adaptive Visual Token Skipping Across Layers For Efficient MLLM Inference
Abstract
Multimodal large language models (MLLMs) require substantial computation to process numerous visual tokens across all transformer layers. Most methods for efficient MLLM inference exploit horizontal redundancy by compressing visual tokens. Beyond token reduction, recent studies exploit vertical redundancy through early exit or fixed-layer skipping. However, we find that the extent and distribution of this redundancy vary across inputs and differ between self-attention and MLP modules. Motivated by these observations, we propose AdaVSkip, which equips each layer with two lightweight routers that independently determine whether visual tokens participate in the computation of the self-attention and MLP modules. These decisions collectively define an input-specific visual computation path. To train these routers, we develop a progressive two-stage training framework that updates only the routers while keeping the backbone frozen. Stage I establishes an initial routing policy by combining supervision from input-specific targets derived from module-wise necessity scores with the downstream language modeling loss. To better account for how routing decisions jointly affect answer quality, we introduce Stage II to further refine the policy. This stage uses reinforcement learning with feedback from answers generated under complete sampled routing paths. It combines an answer correctness reward with a skip-consistency reward that discourages excessive retention of visual token computation. Across three MLLM backbones, AdaVSkip maintains strong task performance with substantially less computation. On LLaVA-NeXT-7B, AdaVSkip reduces language model prefill FLOPs by 53.2% while preserving the original model's average performance. Combining it with visual token compression increases this reduction to 93.1%, while retaining 97.2% of the original performance on average.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.