CV-OPSD: Self-Evolving Long-Context Reasoners via Zero-Annotation Clean-View On-Policy Self-Distillation
Abstract
Models often answer correctly when the relevant clip or span is presented in isolation, yet fail when the same evidence is embedded in a long input. We study this gap as a problem of evidence localization under bounded computation. Our method, clean-view on-policy self-distillation (CV-OPSD), enables a long-context reasoner to self-evolve by pairing a full-input student with a live, stop-gradient teacher instantiated from the same model. The student generates the rollout; the teacher evaluates the same prefixes from a privileged view that reveals where the evidence lies without providing an external answer. As the shared model improves, its clean-view posterior supplies an evolving learning signal without requiring a stronger teacher. In the fully self-generated setting, the examples and their evidence views are produced without human annotation. The view is adapted to each modality: a question-relevant temporal interval for video or marked supporting spans in the identical document. We formalize why such views can increase information accessible to a bounded predictor even when they do not add answer information. Across 9B and 4B backbones, CV-OPSD improves performance across video and long-text evaluations, with gains of up to 4.8 and 5.0 points, respectively. Controlled comparisons with hard SFT, same-view voting, and alternative teacher inputs show that the useful signal comes from transferring an evidence-conditioned posterior on states visited by the full-input policy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.