acceptodds
Under review as a conference paper at ICLR 2027

The Statistical Resolution of Node-Classification Benchmarks

Abstract

Prediction errors of a message-passing network are correlated along the edges of the graph it is evaluated on, so the nodes of a node-classification test set do not contribute independent observations. We study the consequence: the statistical resolution of a benchmark, the smallest difference it can detect, is far coarser than its test-set size suggests, and most published differences fall below it. We formalize resolution through the design effect and the minimum resolvable gap, show that the classical plug-in estimator of is unidentified on real error fields, where a previously undescribed degree-sampling bias alone flips the sign of its correlation estimate on 14 of 16 configurations on ogbn-arxiv. We instead estimate with a graph-partition block bootstrap validated against known design effects, then train 22,400 models on 14 benchmarks. Resolution attaches to the (dataset, model class) pair: on identical Cora test nodes a graph-free MLP has and a GCN . Depth leaves it flat, and size does not buy it, since the 48,603 test nodes of ogbn-arxiv carry the information of roughly 700. Nothing in the bootstrap needs the statistic to be a mean, so the same machinery gives the design effect of ROC-AUC; auditing all 86 published gaps on the scale each was reported on, 87% fall below their benchmark's measured resolution. A claim about an architecture needs one term more: training randomness is a median 28% of the variance of a paired comparison, and adding it takes the rate at which a seed change alone is called significant from 11.9% to 0.4%. Code, the cached per-node correctness vectors and the audit corpus accompany this submission.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.