PairShift: A Controlled Benchmark for Graph Foundation Models under Density Shift
Abstract
Evaluating the structural robustness of graph foundation models (GFMs) often yields a paradox: poorly performing models can appear deceptively robust, whereas highly capable models often prove brittle. Existing cross-dataset evaluations fail to isolate this failure mode, as topology, features, labels, and sampling confound one another. In this paper, we introduce PAIRSHIFT, a benchmark of paired density views: nested edge sets that share nodes, features, labels, and edge randomness. Its context-by-query response surfaces and an exact decomposition distinguish ordinary density effects from a matched-density interaction. Across six GFMs, G2T-FM and GraphPFN incur 16–17 percentage points of mean transfer failure on the common grid, rising to 19 points for G2T-FM on an unseen universe; the other four systems show much smaller interactions, sometimes alongside weak accuracy or unstable predictions. Fixed-checkpoint interventions then show that GraphPFN’s graph-convolution adapters increase mean transfer failure by 11–44 points across controlled synthetic settings, with effects that depend on shift direction and adapter depth. Motivated by this pathway, we construct a class-compatibility neighbor residual with no learned parameters. It improves G2T-FM and SAMGPT across two evaluated universes (2–3 and 1–2 points, respectively) and GOBLIN by 4.4 points in a post-hoc extension, but harms models when additional neighbor evidence is redundant or unreliable. A local error-alignment analysis predicts the direction of its cross-entropy change in all 24 decisive tested conditions. PAIRSHIFT thus connects controlled structural evaluation to a causal failure pathway and a bounded, mechanism-inspired correction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.