Do MLLMs Exhibit Structured Perception Process?
Abstract
While mechanistic interpretability has yielded profound insights into the inner workings of Large Language Models (LLMs), the internal mechanisms of Multimodal Large Language Models (MLLMs) remain largely unclear, particularly regarding how they retrieve and organize visual evidence for complex visual understanding. In this work, we elucidate the mechanisms of visual understanding in MLLMs by analyzing their internal attention dynamics. We find that MLLMs exhibit two distinct patterns of perception (dubbed semantic and reasoning perception) which are concentrated in the middle and later layers of the model. Specifically, **semantic perception** attends to semantic targets of the query, showing highly visual–language alignment. **Reasoning perception** retrieves answer-critical visual evidence at query-boundary tokens. These findings indicate that MLLM visual understanding is not a single-step process but a structured perceptual process. We further demonstrate the functional necessity of these patterns: suppressing the identified heads results in distinct error profiles. Motivated by this insight, we propose Structured Attention Focusing (SAF), a training-free method that emphasizes semantic and reasoning regions via attention-guided enhancement. Across four MLLM backbones and eight benchmarks, SAF consistently improves task scores, including an average gain of 3.3% on Qwen2-VL and 4.1% on Gemma. Our work provides critical insights into MLLM's internal visual mechanisms and introduces an interpretable strategy for enhancing multimodal visual understanding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.