Trust Region On-Policy Distillation
Abstract
On-Policy Distillation (OPD) is a fundamental technique for efficient post-training of large language models (LLMs), with broad applications in agent learning, multi-task enhancement, and model compression. However, OPD training becomes unstable when the teacher and student distributions differ substantially: teacher supervision on student-generated tokens may then yield unreliable policy gradients and even cause optimization failure. To make token-level on-policy supervision reliable, we propose Trust Region On-Policy Distillation (TrOPD), which consists of three components: 1) Trust-Region On-Policy Learning: TrOPD performs OPD only in regions where the teacher provides reliable supervision, mitigating the optimization difficulty of the K_1 reverse-KL estimator under distribution mismatch. 2) Outlier Estimation: For outliers outside the trust region, we compare clipping, masking, and forward-KL estimation, and adopt top-k forward KL to retain informative supervision without unreliable gradients. 3) The student imitates teacher-generated prefixes via forward KL and then continues generation on-policy, steering exploration toward the trust region. Experiments show that TrOPD consistently outperforms SoTA OPD baselines, including OPD, EOPD, and REOPOLD, across mathematical reasoning, code generation, instruction following, and STEM benchmarks. Our code is publicly available to further benchmark advanced OPD methods in community.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.