Distilling Interaction Spaces for Graph Transformers
Abstract
Graph Transformers (GTs), with sophisticated attention mechanisms, naturally support beneficial distant node interactions and have shown strong potential across graph applications. Yet, in semi-supervised node classification, their flexibility makes the choice of interaction structure particularly important: a single interaction graph can over-constrain the student, whereas unrestricted structure learning can be weakly constrained by scarce labels and co-adapt with attention. To bridge this gap, by a constrained bi-level optimization reformulation, we disentangle graph searching and attention learning to reduce detrimental structure–attention co-adaptation and improve generalizability. Furthermore, inspired by weak-to-strong distillation, we distill a task-informed interaction space from a structurally simpler GNN teacher and let the more expressive GT student adaptively search for a suitable candidate within it. This way not only flexibly encodes task-specific implicit constraints on graphs tailored for downstream tasks, but also incorporates students’ model-adaptiveness beyond hand-crafted explicit heuristics. Technically, we propose a two-stage framework: 1) the Disentangling Stage extracts several diverse and task-specific prototypical graphs as a graph basis that spans a reasonable interaction space, with their task relevance enhanced by our Multi-Aspect Cooperative Distillation; 2) the Fusing Stage conducts adaptive search in this space via a simple node-wise linear fuser trained with our Meta Self-Update mechanism. Our framework fully mines the value of teacher knowledge and better releases students’ potential. It empirically outperforms state-of-the-art counterparts while retaining generality and practical efficiency.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.