Beyond a Single Representation: Steering Shared and Modality-Specific Information in Multimodal SSL
Abstract
Multimodal self-supervised learning aims to learn general representations by ex- ploiting interactions between multiple sources of information. However, most existing approaches learn a single representation with a fixed balance between information shared across modalities and information specific to each modality. We argue that such a fixed operating point might be a limitation for multimodal foundation models, where future downstream tasks are unknown during pretraining and may require different information profiles. Hence, we introduce STEER, a preference-conditioned multimodal learning framework that parameterizes a family of representations within a single model. STEER controls the balance between shared and modality-specific information through a preference vector defined on a simplex, and uses low-rank preference-conditioned adaptations to continuously steer the representations toward different information regimes. This allows the de- sired operating point to be automatically selected after pretraining according to the downstream task, without retraining the representation model. We evaluate STEER in both controlled and real-world multimodal settings. In the former, we show that the learned preference space recovers the expected shared and unique information structure and expands the attainable Pareto region beyond fixed-representation baselines. In the latter, STEER outperforms fixed-representation baselines on sev- eral tasks while selecting distinct operating points from the same pretrained model. Importantly, the selected preferences also provide an interpretable description of the balance between modalities favored by each task, which is particularly valuable in scientific applications such as neuroimaging.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.