MIDAS: Multi-Teacher Knowledge Distillation with Adaptive Structure-Aware Embeddings
Abstract
Graph Neural Networks achieve strong performance on graph-structured data but incur high inference latency due to iterative neighborhood aggregation. Knowledge distillation into lightweight multi-layer perceptrons (MLPs) addresses this issue. However, these methods are constrained by the choice of a single teacher, which can be sub-optimal in structurally diverse regions. We propose MIDAS, a multi-teacher distillation framework designed to obtain optimal supervision from GNN architectures across structurally diverse regions, with a student MLP that benefits from the locally optimal teacher rather than a globally best one. MIDAS realizes this through: (1) a dual-encoder contrastive alignment module that embeds structural identity into node representations without graph access at inference, (2) a density-based clustering and cluster-conditional teacher routing mechanism that assigns each node to its locally optimal teacher, and (3) a trunk-adapter student architecture with a two-phase training strategy that prevents gradient interference across divergent teacher signals. Experiments on seven benchmark datasets show that MIDAS outperforms 13 baselines across single-teacher and ensemble distillation settings, achieving up to 80 inference speedup
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.