acceptodds
Under review as a conference paper at ICLR 2027

OWLET: On-Policy Weighted Learning via Evidence Transfer for Video Large Language Model Post-Training

Abstract

Post-training has become a key stage for eliciting reasoning capabilities in Video Large Language Models (Video-LLMs). Existing post-training methods mainly follow two paradigms: reinforcement learning (RL) and on-policy distillation (OPD). However, they do not fully explore visual evidence for effective training. RL supervises the policy with a single trajectory-level reward, which is too coarse-grained for video reasoning: the same response-level advantage is applied to every generated token, disregarding that the tokens of a video response are grounded in the visual stream to different degrees. OPD provides token-level supervision through output-distribution matching, but lacks direct supervision of the spatiotemporal representations that produce those outputs. Our key observation is that the same model reasons more accurately when its video view combines global temporal context with question-relevant details. Comparing the same student-generated response under this view and uniform sampling reveals changes in visual token and internal representations, providing a training signal for evidence use. Building on this observation, we propose OWLET, an evidence-privileged RL method that derives training signals from a richer teacher view. This view couples dense global coverage with question-conditioned local windows, providing additional visual evidence without requiring a larger model. Specifically, OWLETintroduces (i) evidence-aware credit redistribution, which compares token likelihoods across the two views and redistributing positive response-level advantages according to their relative evidence support; and (ii) spatiotemporal state alignment, which aligns the two views' response-end hidden states on verifier-accepted responses. To facilitate training at scale, we release Video-Evidence-96K, containing 95,705 automatically curated question-conditioned temporal-evidence records over 21,832 videos. Experiments on the Qwen3.5 and Qwen3-VL families, ranging from 0.8B to 27B parameters, demonstrate the effectiveness of OWLET, yielding up to 10.9 accuracy improvement compared to vanilla model.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.