Dynamic Multimodal Graph Learning with Content-Conditioned Edge Reweighting and Hierarchical Alignment
Abstract
Multimodal graph learning combines relational structure with textual and visual attributes, but the utility of a neighbor depends on the target content. We introduce DMGraph, which learns content-conditioned weights over induced multimodal subgraphs and aligns text and vision at node, subgraph, and batch-context levels. Its dynamic edge scorer learns rank-controlled bilinear relevance functions, while hardness-aware contrastive learning handles semantically similar negatives. Across WikiWeb2M, FB15K237, Ele-Fashion, and Goodreads-LP, DMGraph consistently improves over graph-multimodal baselines. In expanded neighborhoods, intermediate learned sparsity consistently outperforms dense reweighting and hard Top-K. Under 50% edge corruption, DMGraph limits relative performance degradation to about 15% on both evaluated datasets, demonstrating effective selection of task-relevant graph context.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.