acceptodds
Under review as a conference paper at ICLR 2027

WorldAlign: Decoupled 4D Reward for World-Consistent Video Generation

Abstract

Faithful visual world simulation requires generated videos to maintain 4D world consistency: static environments should preserve coherent 3D structure across viewpoints (static consistency), while dynamic subjects should exhibit plausible motion and consistent appearance over time (dynamic consistency). Existing approaches either modify the generator architecture to incorporate explicit geometry or improve pretrained models through geometry-aware post-training. The former may compromise the broad generalization acquired through large-scale pretraining. The latter preserves this valuable generalization capability but remains unreliable and limited in dynamic scenes. These methods either assume fully static scenes or accommodate dynamic scenes but risk conflating legitimate subject motion with undesired distortions in static regions when evaluating static consistency. Dynamic consistency is either overlooked or evaluated solely through appearance preservation. These limitations prevent effective improvements in world consistency in real-world dynamic scenes. In this paper, we introduce WorldAlign, a decoupled 4D reward framework that semantically separates static regions and dynamic subjects and aligns each with a world prior suited to its assumptions. For static regions, WorldAlign aligns static geometry with a geometric world prior through masked reprojection, excluding dynamic subjects while retaining undesired distortions in static regions for evaluation; an auxiliary camera-motion reward discourages nearly static solutions. For dynamic subjects, WorldAlign uses a strong vision-language model as a dynamic world prior, with sample-specific checklists assessing dynamicity, physical plausibility, shape, and texture consistency. This decoupled design enables online post-training without modifying the generator architecture or requiring human preference annotations. Across two pretrained image-to-video generators, WorldAlign jointly improves static and dynamic consistency over existing methods without suppressing overall or subject motion. These results demonstrate the effectiveness of decoupled world-prior alignment for advancing video generation toward faithful visual world simulation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.