acceptodds
Under review as a conference paper at ICLR 2027

Deep Message-Passing GNNs Have a Backward Problem: Kronecker-Aligned Feedback Training

Abstract

Efforts to train deep message-passing graph neural networks have largely focused on modifying the forward computation to mitigate oversmoothing. However, preserving informative hidden representations alone does not ensure effective learning: supervision can still attenuate through the backward chain of graph propagation, feature transformations, and activation gates. In plain ten-layer GCN and GraphSAGE models on three citation graphs, backpropagation (BP) produces hidden-layer gradients too small to change any FP32 parameter, while frozen probes on the same hidden states remain well above chance. This motivates treating backward credit assignment as an explicit design choice. We introduce Kronecker-Aligned Feedback Training (KAFT), which exploits the Kronecker structure of graph-convolution Jacobians to deliver graph-aware feedback directly to hidden layers, leaving the forward and inference computation unchanged. KAFT combines a capped operator on the forward graph with learned feature-feedback matrices, aligned to directional targets obtained by propagating and normalizing probes through downstream weights. We show that this target satisfies a positive-semidefinite alignment property and give sufficient conditions under which the KAFT update descends and converges to stationarity. Across four node-classification datasets, four backbones, two depths, and three learning rates, KAFT improves accuracy over BP by 3.1 percentage points on average and is significantly better in 76 of 96 matched settings and significantly worse in 4; gains shrink on backbones with built-in identity paths such as GIN. The benefit grows once depth exceeds the point where BP stops updating: on DBLP it rises from 4.3 points at 8 layers to 30.4 points at 20 layers. With BatchNorm or LayerNorm in the same forward model, KAFT is significantly better in four of six comparisons and never significantly worse. In preliminary three-seed experiments, KAFT raises Cora link-prediction ROC-AUC from 64.5% to 81.4% with a six-layer encoder and matches BP on mini-batch graph classification. These results suggest that the backward rule, not only the forward operator, is a useful design axis for deep message-passing GNNs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.