IG-OPD: Information-Geometric Control for On-Policy Distillation
Abstract
On-policy distillation (OPD) trains student language models using teacher feedback on self-generated trajectories, but such feedback can become unreliable when student-generated contexts deviate from teacher-preferred trajectories. In particular, the teacher distribution conditioned on a deviated student prefix may induce a correction direction that conflicts with the target associated with a selected task-valid reference prefix. We study this phenomenon from an information-geometric perspective by modeling next-token distributions on the probability simplex equipped with the Fisher metric. Using the square-root embedding of the categorical Fisher simplex, we derive an exact finite-step identity and a necessary-and-sufficient condition under which an idealized Fisher–Rao correction toward the teacher target conditioned on a deviated prefix moves away from a selected reference target. We further exploit the relationship between Fisher–Rao distance and the Bhattacharyya coefficient to derive a tractable teacher–student discrepancy signal. Guided by these geometry-motivated heuristics, we propose IG-OPD, which combines geometry-aware token weighting, discrepancy-guided prompt routing between off-policy supervised fine-tuning and OPD, and token-wise adaptive interpolation between forward and reverse KL objectives. Experiments across textual and multimodal tasks, teacher configurations, and student scales consistently improve over strong OPD baselines. Controlled-prefix diagnostics further demonstrate that larger teacher–student discrepancy is associated with more frequent reference-direction conflict.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.