acceptodds
Under review as a conference paper at ICLR 2027

When Modality Goes Missing: The Spurious Alignment Trap in Multimodal Large Language Models

Abstract

Multimodal Large Language Models (MLLMs) are generally assumed to degrade predictably when visual input is missing or corrupted, yet existing evaluations focus mainly on downstream performance drops, leaving the internal geometric degradation of representation spaces largely unexplored. In this work, we conduct a subspace geometric analysis of MLLM inference under visually degraded inputs. By measuring the effective rank and principal angles between vision and text subspaces across Transformer layers, we reveal a phenomenon termed the **Spurious Alignment Trap**: when visual input degrades to a blank image, the vision subspace loses dimensionality sharply, while conventional angle-based alignment metrics may counterintuitively indicate improved cross-modal alignment. To distinguish genuine alignment from this illusion, we introduce Gaussian noise images as a diagnostic control. This contrast suggests that the angles under blank conditions are artifacts of dimensional collapse. More importantly, further analysis reveals that under collapse, such metrics decouple from cross-modal semantic alignment and instead reflect the model's modality coupling structure and the landing direction of collapse. Building on these observations, we propose the **Rank–Angle Diagnostic (RAD)** to detect whether a sample's visual input has degraded during inference. Experiments on LLaVA-1.5 and Qwen2.5-VL across VQA and Image Captioning show that blank inputs consistently induce vision-subspace collapse and suppress the expansion of effective rank in deeper layers; across three types of degradation (15%–50% area masking, Gaussian blur, and complete degradation), RAD reliably detects complete degradation, while its sensitivity to partial degradation varies across architectures. Experiments on SmolVLM further validate the coupling–direction relationship on a third model. Our work reveals when and why angular alignment metrics become unreliable, while RAD-based detection requires neither labels nor downstream evaluation. Code will be released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.