acceptodds
Under review as a conference paper at ICLR 2027

Corpus-as-Model Pretraining for Zero-Shot Knowledge Graph Reasoning

Abstract

Recently, knowledge graph (KG) reasoning has shifted from graph-specific inference toward zero-shot knowledge graph foundation models (KGFMs) that transfer relational priors across graphs. Yet prior progress has been driven mainly by architectural advances, while the pretraining corpus has remained a passive artifact, making it unclear whether models learn transferable reasoning patterns or partly adapt to the statistics of a few source graphs. CAMP-KG studies corpus-as-model pretraining: it asks whether a deliberately designed synthetic pretraining pipeline can replace real-KG pretraining while keeping the zero-shot reasoner unchanged. It (i) treats the pretraining distribution as a design object, (ii) samples relational rules independently of graph statistics, and (iii) uses provenance-aware supervision that keeps rule premises observable and prevents target leakage. On the 54 zeroshot graphs of the standard ULTRA suite, CAMP-KG is comparable to real-KG pretraining in aggregate, improves the 32 graphs that share no provenance with the real pretraining sources, and regresses on the 22 source-derived graphs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.