Effective Depth: A Scalar Correction Closes the Oversmoothing Gap in Trained GNNs
Abstract
Asymptotic analyses of oversmoothing in graph neural networks assume linear dynamics, i.i.d. layer weights, and constant aggregation across depth. Trained networks violate all three, and these violations are widely cited as the reason the theory fails in practice. Across 324 training runs spanning three architectures, six depths (2–64), and homophilous and heterophilous benchmarks, we ask which violation matters. The i.i.d. assumption is violated severely – the top singular value of trained weights rises from 0.97× to 2.15× the Marchenko–Pastur bulk edge, pooled across architectures – yet trained networks sit no further from the linear i.i.d. prediction than randomly initialized ones (mean absolute slope error 0.071 against 0.188). What explains the gap is one scalar: the total branch gain∑︁ ℓ ∥hℓ∥/∥xℓ∥ in the residual stream, the network’s effective depth. Training collapses it from 49% of nominal depth at L = 8 to 14% at L = 64, following Leff ≈ cLk with k = 0.39±0.04. Supplying that scalar to an otherwise unchanged linear i.i.d. simulation cuts prediction error by 83% at depth 32 and 89% at depth 64. Its layerwise distribution carries no additional information: no profile of equal total predicts worse than the measured one – uniform, reversed, sorted, and three independent shuffles all fall within 0.80–1.08× (p > 0.18) – while halving the total more than doubles the error (2.16×, p < 10−5). Fitting the exponent on three depths predicts a held-out depth for the same dataset and architecture at about 1.1 times the oracle’s error. Finally we intervene: regularizing total gain in 32-layer GCNs sets it precisely (r = 0.992) and moves decay as predicted (r = −0.943), with unregularized networks on the same curve. Where the constraint binds, decay varies several-fold at no measurable accuracy cost on Cora and Chameleon; Texas, the smallest graph, pays a real cost, and at extreme targets accuracy tracks distance from a network’s own operating point at least as well as it tracks decay. Primary claims are established for GCN, the one architecture where propagation operator and energy metric coincide
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.