acceptodds
Under review as a conference paper at ICLR 2027

RGB-OPD: Representation-Guided Budgeting for On-Policy Distillation in Multimodal Reasoning

Abstract

On-policy distillation (OPD) transfers reasoning capabilities from a teacher to a student on student-generated trajectories, providing supervision at the states actually visited during training. However, we find that image-induced changes are distributed much more broadly in teacher representations than in sampled-token output probabilities, leaving many representation-sensitive positions under-emphasized by standard OPD in multimodal reasoning. Motivated by this representation–output mismatch, we propose RGB-OPD, a representation-guided budgeting framework for multimodal on-policy distillation. RGB-OPD uses teacher representation sensitivity to rank response positions, keeping distillation entirely in the output space. It applies coarse distributional supervision over a compact support constructed from both teacher and student distributions, and reallocates a fixed supervision budget toward the most representation-sensitive positions. A bounded termination correction further addresses truncated reasoning trajectories. Across five out-of-domain multimodal reasoning benchmarks, RGB-OPD improves Standard OPD by 0.80 percentage points in the Qwen3-VL 8B2B run and by 0.91 percentage points in the 32B8B run, with higher point estimates on all five benchmarks at both scales.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.