HumanPhys: A Human-Centric Physical-World Perception Benchmark for Real-World Visual Reasoning
Abstract
We introduce HumanPhys, a large-scale benchmark and training corpus for evaluating and improving human-centric physical-world perception in multimodal large language models (MLLMs). While recent multimodal systems have made rapid progress on visual understanding, they still struggle with basic physical reasoning about humans in 3D scenes, such as spatial relations, motion, and deformation. Existing benchmarks capture only a limited subset of these abilities. HumanPhys is designed to systematically probe six dimensions of human-centric physical perception: spatial state, kinematics, behavior, deformation, light-shadow reasoning, and camera-pose trajectory. These dimensions are instantiated as thirteen subtasks: depth estimation, distance estimation, displacement estimation, speed estimation, trajectory understanding, gaze estimation, person counting, blink counting, facial deformation recognition, body-shadow judgment, facial-shadow judgment, environment-centric camera-pose estimation, and human-centric camera-pose estimation. To support both evaluation and model post-training, we further construct an accompanying training corpus of approximately 760K samples. A key challenge in building such a benchmark is obtaining supervision that is both scalable and physically grounded. To this end, we combine curated real-world datasets—JRDB, UCSD, DanceTrack, Eyeblink8, VideoAttentionTarget, and Decaf—with large-scale simulation in Blender with ReplicaCAD. This hybrid pipeline enables clean supervision over diverse perceptual phenomena that are difficult to annotate consistently in the wild. We benchmark a range of open- and closed-source MLLMs on HumanPhys and find that the benchmark remains challenging across all six capability dimensions. We hope HumanPhys will support future research on multimodal systems for human-centered real-world perception and interaction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.