Mechanistic Signs of Modality Collapse in Multimodal Event Understanding
Abstract
Modality collapse occurs when a multimodal model relies on one dominant modality and ignores others. So far, this phenomenon has been established behaviorally, by modifying the input prompt and analyzing the model's predictions. Such analyses cannot show *whether* a model's computational flow actually leveraged information in a modality. In this work, we focus on the task of identifying important moments in football games and mechanistically verify modality collapse through causal intervention. We separately ablate a moment's video and its commentary transcription from the model's internal representations rather than the prompt, and quantify the impact each modality has on the final prediction. Using the signs of these impacts, we place each moment in one of four diagnostic quadrants: synergy, video collapse, commentary collapse, or conflict. Collapse takes opposite directions across the two classes: where it occurs, important moments are decided by the video and non-important ones by the commentary. Even when a model classifies a moment correctly, collapse costs it confidence. Synergy occurs predominantly for important moments and conflict for non-important ones. Within the important class, synergy monotonically decays from prototypically important moments (goals) to contextually important ones (shots-on-target). To complement these intervention-based results, we also observe how each modality's attention share and the model's decision evolve across layers. Together, our findings indicate that collapse is as much a characteristic of the moment as of the model. Future work can build on these insights to causally localize *where* collapse occurs for enabling inference-time mitigation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.