acceptodds
Under review as a conference paper at ICLR 2027

Video Reversal Curse: Diagnosing and Aligning Temporal Reasoning in Diffusion VLMs

Abstract

Large Language Models can exhibit the Reversal Curse: after learning that event A precedes event B, they may fail to infer that event B follows event A. We show that video-language models exhibit the analogous Video Reversal Curse, combining directional bias with weak reliance on temporal order: they answer cause-to-effect questions more reliably than paired effect-to-cause questions and may ignore frame reversals or shuffles. Although dVLMs provide bidirectional attention and iterative denoising, these capabilities alone do not ensure symmetric temporal reasoning. We propose ReVID, which combines Temporal-Symmetric GRPO, training on paired forward/backward questions with rewards for balanced accuracy and temporal reliance, and Temporal Visual Contrastive Decoding, which contrasts original and shuffled-video logits at inference. We introduce RevVideo, a 3,334-question paired benchmark, and the Temporal Symmetry Score. ReVID improves RevVideo accuracy by 8.37 points over its backbone, reduces the forward/backward gap to 1.78 points, and achieves the best TSS of 40.56, while improving dVLM performance on TempCompass and TVBench.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.