acceptodds
Under review as a conference paper at ICLR 2027

Teaching What the Student Lacks: Reasoning-Mode Transport for On-Policy Distillation

Abstract

On-policy distillation (OPD) trains a student model on its own reasoning trajectories with feedback from a stronger teacher model. However, OPD does not explicitly address imbalanced coverage of reasoning modes at student-visited states. We introduce Reasoning-Mode Transport On-Policy Distillation (RMT-OPD) to prioritize under-covered teacher modes. At states visited along student trajectories that fail final-answer verification, RMT-OPD samples multiple next steps from the student and an answer-conditioned teacher. It uses shared semantic embeddings and optimal transport to compare their local distributions and assign mode deficits. These deficits weight teacher steps during maximum-likelihood training, while subsequent states remain student-generated. RMT-OPD requires only teacher-generated text, without access to teacher logits. Under an idealized discrete-mode model, the weighted deficit equals the teacher probability mass missing from the student. On four mathematical reasoning benchmarks, RMT-OPD performs best among the compared methods and improves Avg4 by 1.7 percentage points over uniform weighting under matched per-state sampling. Fixed-state analyses show smaller teacher–student mode-coverage gaps for ours than MOPD, while gains persist with a larger teacher and on code reasoning tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.