acceptodds
Under review as a conference paper at ICLR 2027

VideoMend: Mending Temporal Continuity in Video Reasoning

Abstract

Video reasoning requires models to capture fine-grained temporal dynamics, includ-ing action directionality, motion dynamics, and state transitions. Yet, constrained by computational costs, many Video Large Language Models resort to sparse frame sampling, which fragments the temporal stream into isolated visual snapshots and encourages single-frame bias: models derive answers from salient objects in individual frames rather than from the temporal evolution of events. We propose VIDEOMEND, a framework that explicitly reasons over inter-frame changes to mend the temporal discontinuity introduced by frame sampling. Its core is Ref-Delta Chain-of-Thought (Ref-Delta CoT), a structured reasoning format organized around two complementary signals: Reference anchors the static scene semantics at temporal anchor points, while Delta explicitly captures the changes between adjacent frames, encouraging the model to articulate temporal transitions rather than derive answers from individual frames in isolation. To train it, we adopt a two-stage pipeline. The SFT stage provides cold-start initialization with data specifically targeting temporal direction and speed perception, including antony-mous semantic pairs that supervise temporal direction and symmetric constructions that decouple semantic complexity from speed. The RL stage further enforces the visual faithfulness of Deltas via Delta-Mask GRPO (DM-GRPO), which constructs counterfactual inputs via delta frame masking and uses a judge model to provide an explicit training signal for the visual grounding of Deltas, ensuring inter-frame changes are grounded in visual observation rather than language priors. Experiments on six benchmarks show consistent gains over strong baselines, most notably on TempCompass, reflecting VIDEOMEND’s capacity to capture temporal continuity.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.