BEAST: Boundary-aware Semantic Extrapolation with Debiased Topology for Long-Tailed Text-attributed Graph Learning
Abstract
This paper studies the problem of long-tailed text-attributed graph learning, where severe class imbalance hinders the learning of tail classes due to limited semantic evidence and biased structural contexts. Existing approaches typically augment minority classes to create a balanced graph and train graph neural networks (GNNs) on the augmented data. Despite the progress, their performance is far from satisfactory due to overconcentration among synthesized tail nodes and noisy edges during graph augmentation. Towards this end, we propose a novel framework named Boundary-aware Semantic Extrapolation with Debiased Topology (BEAST) for long-tailed text-attributed graph learning. The core of our BEAST is to generate informative nodes along with reliable edges in augmented graphs from two complementary perspectives, i.e., semantic boundary exploration and structural pattern acquisition. In particular, we first identify semantic class competitors of each tail class by measuring inter-class prototype distance in the latent space. More importantly, we leverage large language models (LLMs) to generate auxiliary subclasses for tail classes, guiding node synthesis toward desired regions while maintaining sufficient separation from competing classes. These auxiliary subclasses would ensure diversity during extrapolation, while boundary-aware filtering guarantees semantic consistency of synthesized nodes. In addition, we sample a class-balanced graph as a calibration set to train a shadow model for predicting connectivity patterns, which are subsequently utilized for debiased topology imputation. To mitigate parent-child competition, we further adjust the prediction logits and optimize the model in a multi-task learning framework. Extensive experiments on multiple long-tailed text-attributed graph benchmarks demonstrate the effectiveness of BEAST against representative baselines. Our code is available at https://anonymous.4open.science/r/BEAST-CC22.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.