acceptodds
Under review as a conference paper at ICLR 2027

Is Real-World Training Data Necessary for Generalist Graph Anomaly Detection?

Abstract

Generalist graph anomaly detection (GAD) aims to build a foundation model that detects anomalies on arbitrary unseen graphs without retraining or fine-tuning. Sufficient data are essential for foundation model training, yet generalist GAD still faces a data shortage, as real-world anomalous graphs are scarce and costly to collect and annotate. To fill this gap, we propose AG-FORGE, an Anomalous Graph Generation Forge for automatic synthesis of anomalous graphs, exploring the feasibility of synthetic data-driven training for generalist GAD. Empirically, we find that training on synthetic data achieves performance comparable to training on real-world data, but does not further improve the performance of existing methods due to their limited capacity. To better exploit increasingly large-scale training data, we develop TS-GGAD, a Topology-Semantic Coordinated Generalist Graph Anomaly Detection model that captures complementary topological and semantic anomaly evidence, together with a curriculum learning strategy tailored to large-scale synthetic training. Extensive experiments on 14 real-world datasets demonstrate that TS-GGAD, trained on data generated by AG-FORGE, consistently outperforms state-of-the-art methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.