acceptodds
Under review as a conference paper at ICLR 2027

Teaching the Way, Not the Answer: Privileged Tutoring Distillation for Multimodal Policy Optimization

Abstract

Recent post-training methods, particularly Reinforcement Learning with Verifiable Rewards (RLVR), have significantly enhanced the reasoning ability of Large Vision-Language Models (LVLMs). However, the sparse nature of verifiable rewards provides little token-level supervision for failed rollouts, often leading to inefficient exploration in complex multimodal reasoning tasks. Although policy distillation can offer dense guidance, external teacher based methods introduce substantial computational overhead, while answer conditioned tuning methods potentially expose answer-level information and induce shortcut-like generation behavior. To overcome above limitations, we propose a novel Privileged Tutoring Distillation Policy Optimization (PTD-PO) method for RLVR, whose core idea is to supplement supervision with guidance on the reasoning path rather than the answer itself. Specifically, our PTD-PO first constructs Way-revealing Structured Privileged Hints by considering both spatial attention guidance and intermediate textual reasoning steps, which are further employed to guide the student model toward corrective reasoning paths through in-context-learning. particularly, the student model is still optimized under the original question-only context, while only failed rollouts are aligned with the hint-augmented reference model at the token-distribution level. To further stabilize distillation under the distribution shift between guided and unguided contexts, we introduce a Top-K Jensen-Shannon divergence with tail compensation, focusing on aligning informative tokens while reducing memory overhead. Experiments on LVLMs ranging from 2B to 8B parameters show that our PTD-PO can provide effective dense guidance without relying on answer-revealing supervision and consistently outperforms state-of-the-art counterparts by a clear margin.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.