LeGramJEPA: Negative-Free Multimodal Graph Foundation Models in Latent Space with Volumetric Alignment and SIGReg
Abstract
Self-supervised learning on text-attributed graphs enables graph foundation models that transfer to out-of-distribution graphs without labels. The standard approach aligns a graph encoder and a text encoder with a contrastive objective, pulling positive graph-text pairs together and pushing mismatched pairs apart. This depends on large numbers of negative pairs, and extending it beyond two modalities requires a pairwise alignment term for each pair of modalities, so the number of terms grows quadratically. We present LeGramJEPA, a latent-space approach to multimodal graph self-supervised learning that uses no negatives and whose alignment cost grows linearly in the number of modalities. LeGramJEPA replaces the contrastive term with a volumetric alignment objective and constrains each encoder’s output distribution toward an isotropic Gaussian with SIGReg to prevent embedding collapse. The volume term couples all modalities in a single quantity per anchor, reducing the number of alignment terms from quadratic to linear in modalities compared to cross-modal methods. Pretraining with LeGramJEPA is competitive with contrastive baselines on zero-shot node classification and improves over them on link prediction. Under a three-modality (graph, text, image) training and transfer scheme, LeGramJEPA consistently outperforms its contrastive counterpart across modality pairs
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.