acceptodds
Under review as a conference paper at ICLR 2027

Attend, Transform, or Silence: Operator-Level Visual Skipping for Efficient Multimodal LLM Inference

Abstract

Multimodal large language models (MLLMs) increasingly process long visual-token sequences, making decoder prefill computationally expensive. Existing acceleration methods typically remove visual tokens or skip their updates at the layer level, but these coarse strategies may discard fine-grained evidence or suppress useful operators together with redundant ones. We study visual computation from an answer-observable perspective and find that late visual-token updates can remain large while having little effect on answer representations and generated outputs. By decomposing each Transformer layer into attention and FFN operators, we further show that useful visual computation is operator-dominant and varies across layers. We therefore propose OpSkip, an operator-level visual-token skipping framework that preserves the full visual sequence while selectively bypassing visual attention, FFN, or both. Experiments across four MLLMs and 10 VQA benchmarks reduce computation to 55%–67% of the vanilla TFLOPs while retaining 95.0%–99.5% average performance. OpSkip achieves the highest average retention on three of the four architectures and also provides consistent prefill and time-to-first-token speedups.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.