acceptodds
Under review as a conference paper at ICLR 2027

CausalAlign: Causal Alignment of Video Generation Models for Consistent Ego-Motion

Abstract

Diffusion-based video generation models (VGMs) can produce visually plausible videos, yet their camera ego-motion often drifts over long horizons. Camera ego-motion determines how the viewpoint changes between frames and is essential for maintaining consistent geometry and trajectories. To address this problem, we propose CausalAlign, a causal alignment framework with two components. First, we diagnose how precise ego-motion is represented, used, and retained within VGMs using linear probes and causal interventions, revealing a rise-and-fall pattern during denoising. Second, we construct a delta-token space from a 3D foundation model and align VGM representations at causally identified positions. Delta tokens encode changes between consecutive frames, allowing the alignment to focus on ego-motion rather than static scene content. With Wan2.2 fine-tuned on DL3DV and RealEstate10K, CausalAlign preserves visual quality while improving motion accuracy, geometric consistency, and temporal smoothness. For example, it reduces geometric consistency errors by more than 20% relative to the base model. Moreover, using the fine-tuned VGM, CausalAlign improves downstream world action models (WAMs) on the RoboCasa365 benchmark. Overall, CausalAlign improves ego-motion consistency in diffusion-based VGMs and yields gains in downstream embodied models on motion-intensive tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.