Video Reversal Curse: Diagnosing and Aligning Temporal Reasoning in Diffusion VLMs
Abstract
Large Language Models can exhibit the Reversal Curse: after learning that event A precedes event B, they may fail to infer that event B follows event A. We show that video-language models exhibit the analogous Video Reversal Curse, combining directional bias with weak reliance on temporal order: they answer cause-to-effect questions more reliably than paired effect-to-cause questions and may ignore frame reversals or shuffles. Although dVLMs provide bidirectional attention and iterative denoising, these capabilities alone do not ensure symmetric temporal reasoning. We propose ReVID, which combines Temporal-Symmetric GRPO, training on paired forward/backward questions with rewards for balanced accuracy and temporal reliance, and Temporal Visual Contrastive Decoding, which contrasts original and shuffled-video logits at inference. We introduce RevVideo, a 3,334-question paired benchmark, and the Temporal Symmetry Score. ReVID improves RevVideo accuracy by 8.37 points over its backbone, reduces the forward/backward gap to 1.78 points, and achieves the best TSS of 40.56, while improving dVLM performance on TempCompass and TVBench.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.