Feature Extraction Advantage of Graph Transformers: Primal-Dual Decomposition and Reconstruction-Risk Analysis
Abstract
Graph Transformers (GTs) replace fixed topology-driven aggregation with feature-adaptive attention, yet the source of their feature-extraction advantage remains unclear. In particular, how this adaptivity produces discriminative representations and measurable reconstruction benefits is not understood. This paper develops a unified framework connecting primal–dual decomposition, unrolled sparse coding, and reconstruction-risk analysis. We first prove that linearized GT attention admits an exact finite decomposition: key-induced primal atoms analyze the input, while query-induced dual atoms synthesize the output from the resulting projection coefficients. We extend this viewpoint to masked softmax attention through an exact normalized exponential-kernel expansion. We then show that removing feature dependence from the query–key coordinates reduces attention to a fixed linear graph filter, with GCN and spectral graph convolution arising as topology-driven special cases. This identifies the boundary between adaptive attention and fixed aggregation. Motivated by the decomposition, we interpret a deep GT as an unrolled proximal-gradient solver for an -regularized coding problem: attention generates an adaptive dictionary, and successive layers perform residual correction and group-sparse coefficient refinement. Finally, using a ridge surrogate, we derive explicit lower bounds on the reconstruction-risk advantage of adaptive GT-inspired filters over fixed GCN filters. The shared-basis specialization isolates spectral adaptivity by allowing the two filter families to differ only through their spectral weights. Controlled experiments verify the decompositions and the degeneration to fixed filtering, and show that adaptive attention improves class discriminability and lowers reconstruction risk relative to fixed topology-driven aggregation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.