acceptodds
Under review as a conference paper at ICLR 2027

Tri-JEPA: Interaction-Aware Self-Supervised Multimodal Pretraining

Abstract

Self-supervised multimodal pretraining aims to learn representations from unlabeled multimodal data that transfer across downstream tasks. Different tasks, however, may rely on different multimodal interactions: information shared across modalities, information unique to a single modality, and synergistic information that emerges only when modalities are considered jointly. Since the downstream task is unknown at pretraining time, encouraging all three may yield a more broadly useful prior. Existing methods either emphasize only a subset of these interactions or capture them implicitly through objectives over fused multimodal representations. We introduce Tri-JEPA, an interaction-aware joint-embedding predictive architecture with three predictive objectives defined by different source-target configurations. For a target modality, the shared objective predicts its representation from another modality, the unique objective predicts it from the same modality, and the synergistic objective predicts it from all other modalities jointly. These assignments are cycled across modalities for balanced coverage. The tailored objectives enable us to examine how different multimodal interactions are encouraged during pretraining and how their contributions vary across downstream tasks. Under controlled settings, Tri-JEPA consistently outperforms the multimodal self-supervised baselines considered across diverse downstream tasks. Overall, our results show that interaction-aware predictive objectives are a useful design principle for self-supervised multimodal pretraining.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.