acceptodds
Under review as a conference paper at ICLR 2027

Learn Only What Helps : Task-Prioritized Distillation from Heterogeneous Teachers for 2D Molecular Property Prediction

Abstract

Accurate molecular property prediction (MPP) can benefit from heterogeneous molecular representations that encode complementary structural and chemical information. Yet obtaining such representations often requires additional molecular views or computations that are costly or impractical for large-scale screening. Knowledge distillation (KD) provides a natural way to transfer such supervision to a 2D graph student during training while retaining efficient 2D-only inference. However, existing multi-teacher KD methods largely focus on how to combine supervision from multiple teachers, but overlook whether a given teacher signal helps the downstream objective at a given student update. We argue that this is particularly significant for heterogeneous molecular teachers: differences in gradient magnitude can cause some teachers to dominate the shared update, while directional disagreement can make even a properly scaled teacher interfere with task descent. To address this challenge, we propose First-Order Routing of Teacher Expertise (FORTE), which reframes heterogeneous distillation as a task-prioritized routing problem rather than a teacher aggregation problem. Instead of assigning supervision to all available teachers, FORTE uses the downstream task gradient to determine both whether a teacher update is admissible and how strongly it should influence the student. Compatible teachers are routed through and a null route abstains when none qualifies, yielding an auxiliary update that is both aligned with and bounded by the task gradient.Theoretically, we show that, under the stated assumptions, FORTE improves the convergence bounds for the supervised downstream objective over conventional teacher aggregation, yielding a smaller sufficient iteration budget. Experimental results over a wide range of benchmarks show consistent improvements in both classification and regression, demonstrating the effectiveness and generality of our proposed method.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.