Heterogeneous Graph Prior-Data Fitted Foundation Model
Abstract
Heterogeneous graphs are ubiquitous in real-world applications, yet their schema heterogeneity, relational diversity, and complex attribute distributions make it highly challenging to build foundation models with cross-task generalization. Existing graph foundation models are mostly designed for a single setting, such as homogeneous graphs or relational data, or rely on large-scale real-world pretraining data, and therefore face data scarcity and limited task adaptability in heterogeneous scenarios. To address this, we propose HGraphFM, a foundation model framework for heterogeneous graph node-level tasks, which enables unified modeling of diverse heterogeneous graph tasks through synthetic-data-driven pretraining. In a nutshell, we first design a general heterogeneous graph prior, which, together with the relational-data prior, provides large-scale, diverse pretraining data for the foundation model. HGraphFM extends a single-table model by introducing graph attention layers, which in turn enhance in-context learning attention, enabling a transition from processing single-table data to handling heterogeneous data. It combines both in-context learning and structure learning capabilities, and can jointly model heterogeneous attributes and relations that may be subject to temporal constraints. Furthermore, we introduce a feature augmentation mechanism, termed Deep Feature Synthesis Based on Graph Band-Pass Filtering (DFS-BP), which effectively alleviates insufficient context utilization and the need for manual feature selection. The mechanism, reminiscent of “lazy learning,” produces synthesized features that directly match or even surpass the performance of some existing pretrained models in classification tasks, and provides substantial gains in model generalization. After a two-stage training procedure based on the two types of priors, experiments on multiple real-world heterogeneous graph datasets show that our method outperforms existing approaches, including homogeneous-graph and relational foundation models, in both few-shot and fine-tuned settings, validating the effectiveness of the proposed heterogeneous graph foundation model paradigm.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.