Interactive-Policy Distillation with Bidirectional Propose-and-Verify
Abstract
On-policy distillation (OPD) trains a compact student model on its self-generated trajectories with dense token-level teacher feedback. However, naive OPD may suffer from teacher unanchoring, where the student's reasoning trajectory drifts outside the teacher's acceptance region, causing the teacher to be queried on states it would hardly visit and thus provide unreliable supervision. We propose **Interactive-Policy Distillation (IPD)**, which dynamically corrects training trajectories within adaptive teacher intervention. Under a bidirectional propose-and-verify state machine, the student and teacher alternately exchange their roles as proposer and verifier, and collaboratively generate mixed-source trajectories. Then different supervisions are applied according to the source of each token. This bidirectional propose-and-verify mechanism and the source-split loss make IPD not only a more performant distillation method, but also a unified bridge between on-policy and off-policy paradigms. To make the interleaved dual-model rollouts more efficient, we also design a dedicated fused inference engine that co-hosts both models in one serving instance with separate KV caches and instantiates the state machine model to distribute, collect, and process requests. On math reasoning tasks and across multiple teacher-student model pairs, student models trained with IPD not only outperform those trained with OPD, but also demonstrate higher data efficiency. Specifically, when distilling Qwen3-30B-A3B (non-thinking) into Qwen3-1.7B-Base, IPD brings a mean@8 and a best@8 benchmark-averaged accuracy improvement compared with OPD. Besides, IPD only consumes about of the training examples and steps to outperform OPD trained on the whole training dataset for one epoch. We also investigate the impact of different loss variants and takeover / handback configurations, and demonstrate the robustness of IPD on different training data.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.