acceptodds
Under review as a conference paper at ICLR 2027

Learning to Generalize from the Vicinity: Knowledge Expansion from a Single Teacher

Abstract

When empirical data distributions poorly represent the true data manifold, models exposed to local continuities around data tend to generalize better than those that tightly fit individual data. Similarly, in Knowledge Distillation (KD), when teacher representations poorly reflect the true representation manifold, can exposing a student to local neighborhoods provide greater benefits than fitting to individual teacher targets? In other words, *Can vicinal distributions around teacher targets help students learn better?* We investigate this by introducing a novel distillation objective where we admit a generative vicinal distribution for each individual teacher target; **Generative Vicinity Distillation (GVD)**. GVD learns a generative prior over the teacher representation space to enable sampling from vicinities and derives guidance functions to impose desired constraints on the generated samples (class consistency, locality). Departing from conventional KD thinking that depends on multi-teacher supervision for robustness, diversity, and fidelity, we instead exploit the ambiguity of a generative model’s solution space to derive such properties from a single teacher. Extensive experiments across datasets and architectures demonstrate the effectiveness of GVD in improving student generalization and robustness.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.