Correctness-Conditioned On-Policy Distillation for Reflective Multi-Turn Reasoning
Abstract
Reinforcement-learning-based post-training has improved the reasoning capabilities of vision-language models. Inspired by a teacher–student learning paradigm centered on questioning, receiving correctness feedback, and revising, we propose a Reflective Multi-Turn Video Reasoning method that enables models to verify and correct their own predictions without access to privileged feedback. First, Multi-Agent Dataset Filtering combines diverse reasoning trajectories generated by multiple agents with agent-based evaluation and human verification to construct high-quality supervision. Second, Reflective Multi-Turn Training iteratively retains mixed-outcome samples and jointly optimizes the model using Group Relative Policy Optimization (GRPO) and Correctness-Conditioned On-Policy Distillation (CCOPD), thereby improving answer verification and revision. During the second stage, CCOPD conditions the teacher model on the correctness as privileged information and provides token-level supervision along the student’s on-policy trajectory. We evaluate the proposed method on 13 widely used benchmarks spanning four categories: short-video understanding, long-video understanding, video reasoning, and temporal grounding. The results show that our method achieves broadly competitive performance across all 13 benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.