Would an Untrained Encoder Do as Well? The Missing Baseline in Graph Self-Supervised Learning
Abstract
Graph self-supervised learning is judged by freezing a pretrained encoder, fitting a linear probe, and comparing accuracies on small benchmarks, where methods are separated by a point or two. These comparisons almost never ask whether the same encoder, never trained, would do as well. We run this control for graph joint-embedding predictive architectures (JEPAs) on seven benchmarks, with three partitioners, two read-outs and two implementations, including the official Graph-JEPA code, whose published accuracies we reproduce. We find that pretraining adds no detectable accuracy over a randomly initialised copy of the encoder except on the smallest dataset, and that under the published read-out the untrained encoder is significantly better on three datasets. Pooled over benchmarks, every setting excludes a common gain above one and a half points. Pretraining does reshape the representation, without improving what a linear probe reads, and the most robust partition effect, structural-entropy over METIS partitions on NCI1, is nearly as large for the never-trained encoder as for the trained one. The benchmarks set the boundary: under a corrected test they cannot resolve gains below two to six points, several times the margins that rank published methods. An untrained encoder costs one run with zero epochs, and graph self-supervised learning should be measured against it.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.