Corpus-as-Model Pretraining for Zero-Shot Knowledge Graph Reasoning
Abstract
Recently, knowledge graph (KG) reasoning has shifted from graph-specific inference toward zero-shot knowledge graph foundation models (KGFMs) that transfer relational priors across graphs. Yet prior progress has been driven mainly by architectural advances, while the pretraining corpus has remained a passive artifact, making it unclear whether models learn transferable reasoning patterns or partly adapt to the statistics of a few source graphs. CAMP-KG studies corpus-as-model pretraining: it asks whether a deliberately designed synthetic pretraining pipeline can replace real-KG pretraining while keeping the zero-shot reasoner unchanged. It (i) treats the pretraining distribution as a design object, (ii) samples relational rules independently of graph statistics, and (iii) uses provenance-aware supervision that keeps rule premises observable and prevents target leakage. On the 54 zeroshot graphs of the standard ULTRA suite, CAMP-KG is comparable to real-KG pretraining in aggregate, improves the 32 graphs that share no provenance with the real pretraining sources, and regresses on the 22 source-derived graphs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.