PhysicsLENS: Diagnosing Physical Property Blindness in Video Generation Models
Abstract
Reliable video world models could provide scalable predictive environments for robot learning, planning, and evaluation. However, robot videos can violate physical principles and complete tasks through physically implausible behavior, limiting their reliability for robot learning and planning. Current video-generation benchmarks exclude physics that are inherently hidden by visuals (eg: weight, viscosity, friction). Due to this, video models are evaluated on the fidelity of physics- not the underlying accuracy of physics. We introduce PhysicsLENS, a dataset and benchmark for evaluating plausibility of physical properties grounded in robotics. PhysicsLENS uses matched scenario pairs that hold the conditioning frame, task, and scene description. Scenarios are curated from public robot video sources and annotated across seven physical domains: collision, gravity, momentum, friction, deformation, fluid, and causality. We evaluate across four video generation models, producing over 400 human-annotated labels. Results show a consistent accuracy gap between observable and unobservable conditions, indicating that models rely on visual appearance rather than grounded text-specified properties.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.