acceptodds
Under review as a conference paper at ICLR 2027

Seeing the Evidence: Spatiotemporal-Privileged On-Policy Self-Distillation for Long-Video Understanding

Abstract

Vision-language models still struggle to identify and integrate fine-grained spatiotemporal evidence for long-video question answering. We compare different combinations of temporal and spatial evidence and find that joint spatiotemporal views improve the same model's answer accuracy over standard full-video inputs. Building on this study, we propose ST-OPSD (Spatiotemporal On-Policy Self-Distillation), a framework that transfers the model's own predictive advantage under privileged spatiotemporal views to its full-video policy. Specifically, we construct the privileged visual input by combining temporal evidence frames, bounding-box-marked keyframes, and local crops, preserving event context while highlighting spatial details. A frozen teacher initialized from the same pretrained model as the student provides token-level supervision under this view along student-generated response trajectories. We train ST-OPSD with a forward-KL distillation objective on 3,200 carefully selected examples from CG-Bench, which we annotate with spatial bounding boxes. Experiments show that ST-OPSD achieves strong performance across four video understanding benchmarks using Qwen3.5 models at different scales.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.