EnergyOPD: Energy-Based On-Policy Distillation for Stable and Generalizable Knowledge Transfer
Abstract
On-policy distillation (OPD) has become a mainstream knowledge transfer paradigm for large language models (LLMs), alleviating the exposure bias of off-policy distillation through interactive student rollouts and real-time teacher supervision. However, existing OPD objectives are almost exclusively instantiated as Kullback–Leibler (KL) divergences and related statistical distribution metrics, which suffer from two critical defects. First, single divergence-based losses are highly sensitive to local token noise and trajectory drift, causing severe gradient oscillation and unstable convergence. Second, the rigid distribution-matching framework cannot flexibly integrate heterogeneous supervision signals such as execution outcomes and task rewards, limiting scenario adaptability and out-of-distribution (OOD) generalization. In this paper, we reconstruct on-policy distillation from an energy-based modeling perspective and propose EnergyOPD, a novel energy-based on-policy distillation framework. We model the teacher LLM as a task-oriented energy function and reformulate OPD as fitting the student energy landscape to the teacher energy landscape, deriving a generalized free-energy distillation objective that unifies reverse KL, forward KL, and Jensen–Shannon variants within a single framework. We further propose an energy landscape smoothing regularizer that provably reduces gradient variance and improves OOD generalization, together with a universal multi-signal fusion mechanism that integrates token-level teacher feedback, trajectory execution results, and task rewards into unified energy terms. Experiments on six mathematical reasoning benchmarks across three student scales (Qwen3-0.6B/1.7B/4B-Base), as well as out-of-domain evaluations, show that EnergyOPD consistently improves both average Avg@8 and Pass@8 over strong on-policy distillation baselines, while substantially reducing training instability and achieving higher out-of-distribution accuracy, demonstrating stable, generalizable, and scalable on-policy knowledge transfer.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.