Teacher Imitation Is Not Enough: Exposure-Aware Optimal-Transport Distillation for Recommendation
Abstract
Knowledge distillation (KD) for recommendation tasks typically encourages a compact student to imitate a high-capacity teacher. Yet teachers are trained on logged interactions shaped by uneven historical exposure and popularity feedback, so direct distillation can propagate popularity bias together with useful ranking knowledge. This raises a fundamental question: should a student imitate everything its teacher has learned? We propose EOT-KD, an exposure-aware optimal-transport distillation framework that separates teacher correction from teacher transfer. Using interaction frequency as an exposure proxy, EOT-KD first constructs a minimum-distortion exposure correction of the teacher distribution; the closest distribution in Kullback–Leibler (KL) divergence while penalizing expected historical exposure. This correction admits a closed-form solution and provides monotonic control over the exposure score through a single parameter. The corrected target is then transferred through entropic optimal transport, with pairwise costs defined by cosine distance between teacher-derived item representations, assigning lower transport costs to more similar items. Across four recommendation benchmarks, EOT-KD improves Tail-NDCG@20 over the teacher by 25.9-48.9% while retaining 94.8-99.7% of the teacher's NDCG@20. Teacher and OT computations are used only during training. At inference, only the graph-free student is retained, requiring no teacher, graph propagation, or optimal transport.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.