IP-MoE: Investigating Bilinear Vision–Language Interactions for Expert Weighting
Abstract
Mixture-of-Experts (MoE) architectures provide a mechanism for combining expert computation through input-dependent weighting. In vision–language models, however, it remains unclear whether explicit interactions between visual and textual features provide useful information for determining these weights. We investigate this question with IP-MoE, a soft dense mixture of experts conditioned on token–patch interactions. IP-MoE constructs a spatial interaction map from question-token and image-patch representations while preserving the image layout, and uses it to assign soft weights to global, local, and hybrid experts. Their weighted outputs, together with image and interaction features, are projected into prefix tokens for a causal language model. We evaluate IP-MoE on seven benchmarks spanning visual question answering, spatial perception, object hallucination, general vision–language abilities, and vision-indispensable reasoning. We further compare bilinear and sum-based interactions, as well as learned and uniform expert weighting. In the evaluated configuration, the model trained with learned expert weighting obtains higher scores on 2D spatial perception and object hallucination, but lower scores on 3D spatial perception, than the model trained with uniform weights. Bilinear interactions likewise yield task-dependent differences relative to sum-based interactions. Training diagnostics indicate nearly uniform gates late in training, and restricting inference to individual expert families produces modest changes in aggregate scores. These findings characterize the behavior and limitations of interaction-conditioned expert weighting, without establishing consistent performance gains or functional expert specialization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.