Pitfalls of Evaluating Graph Foundation Models for Node Property Prediction
Abstract
Due to the wide use of graph-structured data in different fields of industry and science, the development of Graph Foundation Models (GFMs) has recently attracted a lot of attention. While many different types of models get called GFMs, particular interest has been paid to GFMs designed for node property prediction tasks, which is one of the most popular settings in Graph Machine Learning with lots of real-world applications from fraud detection in financial and social networks to recommendation systems for e-commerce and user-generated content platforms. While a number of GFMs for node property prediction have been recently proposed, the field has not converged to a unified evaluation setting, and different works evaluate their models in widely different ways. We argue that the currently used evaluation settings often have significant limitations, preventing reliable assessment of GFM quality and comparison of GFMs with each other and to other types of models. In this work, we aim to highlight common pitfalls in the evaluation of GFMs for node property prediction, discuss potential solutions, and draw attention to open questions. We discuss such aspects as dataset and evaluation metric selection, aggregation of results across different datasets, proper tuning of baselines, and computational cost considerations, supporting each issue with illustrative examples from recent works in the field. We conclude with a fair reevaluation of 9 GFMs from the literature, showing that only the most recent GFMs based on the Prior-data Fitted Networks paradigm can compete with (and often outperform) properly tuned classic Graph Neural Networks in predictive performance, although at a much higher inference cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.