ViB-OPD: Visual Evidence Mining and Boundary-Balanced Video Distillation
Abstract
Existing approaches commonly use supervised fine-tuning (SFT) or reinforcement learning (RL) to develop video understanding and reasoning; however, SFT is constrained by the quality and coverage of fixed demonstrations, while RL often relies on sparse final-answer rewards that provide little direct guidance for intermediate reasoning. On-policy distillation (OPD) offers an alternative by providing teacher distributional supervision along trajectories generated by the current student. Yet applying OPD to video reasoning leaves two issues unresolved: teacher distributions can be limited by incomplete visual evidence, and token-wise aggregation can make supervision allocation depend on expression length. We propose ViB-OPD, comprising Visual Evidence Mining (VEM) and Boundary-Balanced OPD (BB-OPD). A stronger frozen teacher uses supplementary training-time visual context to guide the student without expanding its inference input. VEM combines global observation, question-conditioned localization, and optional local reinspection to organize candidate evidence into a separate contact sheet, enriching the teacher's visual basis. BB-OPD initializes semantic groups and combines local boundary perturbation with within-segment averaging and equal segment weights to reduce the influence of expression length and a single segmentation on supervision allocation. It preserves the original student trajectory and token-level distillation objective. Evaluation on six public benchmarks and across model families demonstrates training gains.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.