Semantic Role Projection: Structured Visual Embeddings for Compositional Reasoning in Multimodal LLMs
Abstract
Multimodal large language models (MLLMs) rely on learned projectors to translate visual patch features into the language model's token embedding space. We identify a fundamental limitation of standard projectors— semantic role flattening—whereby object identity, attribute, count, and spatial relation information is collapsed into an undifferentiated embedding region, despite the LLM's own text embeddings maintaining clear role-specific subspace structure. We propose the Semantic Role Projector (SRP), which replaces the standard MLP projector with four parallel typed projection heads supervised by role-alignment and inter-role orthogonality losses derived from the LLM's frozen text-embedding subspaces. A lightweight learned router produces soft role assignments per visual patch, and a mixing layer preserves drop-in compatibility with existing MLLM architectures. We prove that SRP provably increases the attention logit gap for role-selective heads under mild spectral assumptions, providing a mechanistic explanation for its compositional advantage. Experiments on CLEVR-CoGenT, CLOSURE, GQA, TallyQA, and VQAv2 demonstrate that SRP achieves state-of-the-art compositional generalization—improving CoGenT Condition B accuracy by +3.6 pp over the best modular baseline (Honeybee) and +5.3 pp over a capacity-matched Mixture-of-Experts-while adding only 3% wall-clock overhead. Ablations confirm that each geometric loss component is essential, and cross-architecture evaluation on Qwen2.5-VL-7B verifies that the gains transfer across LLM backbones.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.