acceptodds
Under review as a conference paper at ICLR 2027

MechVerse: Evaluating Physical Motion Consistency in Video Generation Models

Abstract

Text- and image-conditioned video generation models have achieved remarkable progress in visual fidelity and temporal coherence. However, they remain fundamentally appearance-driven and fail to model structured motion dependencies between interacting objects. In many real-world systems, motion is interdependent where some components remain stationary while others move, and the behavior of one part constrains or influences others. When these constraints are violated, generated videos may still appear visually plausible but are functionally incorrect, with parts moving independently when they should not and interactions failing to propagate across related elements. To study this limitation systematically, we introduce MechVerse, a large-scale synthetic video dataset of mechanical assembly animations designed specifically for evaluating structured motion understanding in generative video models. MechVerse comprises over 21,156 video clips spanning across kinematic complexity- single-part articulation, two-part coupling, and strongly coupled multi-part mechanisms. Each clip is paired with a structured text prompt that precisely describes part identities, stationary and moving components, motion type, and inter-part dependencies. Using MechVerse, we conduct the first comprehensive benchmark of state-of-the-art text- and image-conditioned video generation models on the task of mechanically-consistent video synthesis, evaluating both open-source and closed-source models. Our results reveal a consistent and significant degradation in generation quality as kinematic coupling complexity increases, demonstrating that current generative video models lack the representational capacity to reason about structured inter-part motion dependencies. MechVerse establishes a rigorous diagnostic benchmark for this underexplored capability and provides a foundation for improving motion-aware generative video models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.