RGB-OPD: Representation-Guided Budgeting for On-Policy Distillation in Multimodal Reasoning
Abstract
On-policy distillation (OPD) transfers reasoning capabilities from a teacher to a student on student-generated trajectories, providing supervision at the states actually visited during training. However, we find that image-induced changes are distributed much more broadly in teacher representations than in sampled-token output probabilities, leaving many representation-sensitive positions under-emphasized by standard OPD in multimodal reasoning. Motivated by this representation–output mismatch, we propose RGB-OPD, a representation-guided budgeting framework for multimodal on-policy distillation. RGB-OPD uses teacher representation sensitivity to rank response positions, keeping distillation entirely in the output space. It applies coarse distributional supervision over a compact support constructed from both teacher and student distributions, and reallocates a fixed supervision budget toward the most representation-sensitive positions. A bounded termination correction further addresses truncated reasoning trajectories. Across five out-of-domain multimodal reasoning benchmarks, RGB-OPD improves Standard OPD by 0.80 percentage points in the Qwen3-VL 8B2B run and by 0.91 percentage points in the 32B8B run, with higher point estimates on all five benchmarks at both scales.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.