PVBench: Quantifying Physical Plausibility in Video Generation with Object-Level Physical Measurements
Abstract
Recent video generation models produce visually compelling outputs, yet their physical plausibility remains poorly measured. Physical plausibility asks whether generated objects obey geometry, kinematics, dynamics, and conservation laws. Existing benchmarks rely on coarse visual-language model judgments or target embodied world models rather than video generation itself. We introduce **PVBench**, **a physical video benchmark that quantitatively measures the plausibility of generated videos across 46 dimensions organized into seven domains**, spanning: *i)* instruction and semantic fidelity, *ii)* behavioral and causal consistency, *iii)* perceptual and temporal quality, *iv)* geometric and spatial consistency, *v)* material properties and field states, *vi)* kinematic and spatiotemporal persistence, *vii)* and mechanical plausibility with conservation laws. **PVBench** comprises 600 carefully crafted prompts, of which 30% involve physical embodied agents such as robot arms, humanoids, quadrupeds, and multi-agent systems. For challenging physically tractable dimensions, it pairs each with specialized vision models that derive verifiable, physically grounded measurements, such as object-level 6D pose, metric scale, surface normals, kinematics, and numerical conservation-law residuals. It thus avoids relying solely on subjective scoring, while supporting both text-to-video and image-to-video generation. Finally, we validate the agreement between automatic metrics and human judgments through human evaluation. Meanwhile, we evaluate mainstream open-source and closed-source models and present several insights. By releasing the benchmark, prompts, and calibration protocol, we aim to foster physically grounded video generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.