Adaptive Alignment-Driven Feature Augmentation for Graph Learning
Abstract
Graph Neural Networks (GNNs) have achieved strong performance on a wide range of graph learning tasks by integrating node attributes with graph structure through message passing. However, when structural and node attributes encode signals that are weak, inconsistent, or misaligned with the supervision objective, message passing can amplify noise and hinder performance. Existing approaches primarily address these challenges through architectural or structural modifications, and existing feature augmentation methods generally do not account for how the downstream GNN propagates the augmented signals. In this work, we adopt a data-centric perspective and propose GraphAlFA (Alignment-Driven Feature Augmentation), an adaptive, propagation-aware graph feature augmentation framework that explicitly enhances feature-structure-label alignment under task supervision. Rather than modifying model architecture or graph topology, GraphAlFA operates directly in the feature space by constructing complementary attribute-derived and structure-derived views, and optimizing them through complementary alignment objectives: supervised contrastive learning promotes label-discriminative representations, while a propagation-aware spectral objective estimates the downstream GNN's empirical frequency response and uses it to emphasize task-relevant spectral components effectively preserved by propagation. By conditioning augmentation on this empirical response instead of a fixed spectral prior, GraphAlFA adapts the input representation to the propagation behavior of the backbone while preserving the graph topology and backbone architecture. Theoretically, we provide an information-theoretic characterization of feature-label alignment and a propagation-aware characterization of spectral feature-structure alignment. Extensive experiments on synthetic and 9 real-world datasets spanning diverse homophily levels show that GraphAlFA consistently outperforms state-of-the-art baselines, achieving average accuracy gains of 4.82% on real datasets and 12.22% on synthetic datasets over original features.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.