Mixture of Primitives: Fine-Grained Expert Sharing for Generative Visual Perception in Unified Multimodal Models
Abstract
Unified multimodal models (UMMs) increasingly support dense visual perception through image generation, an emerging setting that we refer to as Generative Visual Perception (GVP). While GVP tasks share a common generative interface, they require distinct visual transformations to produce task-specific outputs, motivating task-dependent specialization while preserving shared representations. Mixture-of-experts (MoE) provides a natural mechanism for introducing this conditional specialization through routed experts, but conventional MoE parameterization couples routing decisions with expert-level parameter ownership, requiring each expert to maintain an independent copy of its full transformation. Such coupling restricts parameter ownership to the expert level. We propose Mixture of Primitives (MoP), a fine-grained expert-sharing mechanism that decouples routing identity from parameter ownership by selectively sharing internal channel-group primitives across experts, enabling fine-grained parameter sharing while preserving route-specific transformations. We mainly validate MoP on Cheers, a pretrained UMM, forming Vision-Cheers, which achieves competitive performance with task-specific models while retaining image generation and multimodal understanding capabilities, and further validate MoP on Bagel for cross-architecture generality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.