Toward Omni Multimodal Graph Foundation Model: A Topology-Driven Binding Approach
Abstract
Multimodal graph foundation models (MGFMs) seek to learn generalizable representations from large-scale graphs with heterogeneous node modalities. However, real-world Multimodal-Attributed Graphs (MAGs) often contain incomplete node attributes, limiting the scale and diversity of available pretraining corpora. Besides, existing MGFMs primarily incorporate graph topology as structural context, overlooking its role in guiding multimodal binding and shaping a unified representation space. To address these challenges, we propose , a topology-driven approach that uses graph topology to bind rich modality information into a unified shared space. GraphBind is motivated by the stability of graph topology, which provides structural references and complementary semantic information for multimodal binding. Concretely, GraphBind uses topology to organize self semantics and reliable neighborhood semantics into a global shared space that integrates structure and semantics, and adapts this space to discriminative and generative tasks through lightweight interfaces. Extensive experiments against 11 representative baselines demonstrate that GraphBind achieves leading performance on both discriminative and generative tasks, with relative improvements of up to 28.1% over the strongest baseline.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.