Before Attention: Functional Visual Token Roles at the Vision–Language Interface
Abstract
Large vision–language models connect visual encoders to pretrained language models through a vision–language interface, yet visual-token function is typically characterized only after decoder processing begins. We ask whether functionally predictive differences among visual tokens are already accessible at this interface. Across projector- and merger-based LVLMs, we observe a recurring post-interface organization: activation mass becomes less concentrated within individual tokens, while a restricted set of hidden coordinates is repeatedly reused across tokens. To test whether this structure is functionally meaningful, we develop PRISM, a lightweight role-recovery procedure that uses paired interface-relative concentration responses to construct interface-response anchors, then recovers image-level contributive and residual roles from dominant coordinate identities and relative ranks. Controlled interventions reveal a consistent functional asymmetry: removing contributive tokens is substantially more disruptive than removing residual tokens, even after matching removal count and subsequent decoder attention. The recovered distinction also extends beyond the initial structural contrast, indicating that PRISM captures structure not available from scalar concentration change alone. Because the partition is computed before query-conditioned decoder processing, the same image-level assignment can be reused across questions. These results show that functional differences among visual tokens are already reflected in representations entering the language model, establishing the vision–language interface as an earlier stage for multimodal analysis.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.