acceptodds
Under review as a conference paper at ICLR 2027

CoDance for Physically Consistent Image-to-Video Generation via Coupled Reinforcement Learning

Abstract

Image-to-video models can produce visually realistic videos while violating the physical constraints governing motion and interactions, limiting their utility for embodied prediction and planning. Some recent attempts employ physical reasoning to refine the user prompt before feeding it into the video generation model, making the intended dynamics and interactions more explicit. However, reasoning over the prompt alone does not ensure that the generator can faithfully realize these inferred dynamics. This motivates jointly optimizing reasoning and generation to address the mismatch between inferred physical dynamics and their realization in generated videos. We introduce CoDance with improved physical quality by jointly optimizing a vision-language reasoner and a flow-based video generator with reinforcement learning. The resulting training pipeline, dubbed CoDanceRL, combines three parts: Coupled Policy Optimization, Physics-Aware Feedback, and Hierarchical Reward Normalization. CoDance obtains an IQ-Score of 45.4 on Physics-IQ Verified and an average embodied video generation score of 69.4 on RBench, compared with 34.8 and 58.1 for its Cosmos3-Nano base model. It also achieves an average closed-loop action-planning success rate of 41.5% across two WorldArena tasks, demonstrating its utility for downstream planning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.