SPARKBench: Best Lessons from Benchmarking Human-Centric Spatio-Physics Reasoning
Abstract
Understanding how another person can move and interact with the physical world is essential to human coordination and collaboration. Yet even a simple question—can someone bend down to pick up an object without hitting a nearby table or losing balance?—requires reasoning about articulated motion, collision, and physical stability that remains challenging for multimodal large language models. To study Spatio-Physics in humAn-centric Reasoning, we introduce SPARKBench, a benchmark with 5,291 questions over 1,442 images. It covers two macro-level categories: 3D Spatial reasoning about pose, orientation, and relations, and Spatio-Physics reasoning about collision, articulation and kinematics, and intuitive physics. We benchmark 12 vision-language models, finding that all of these models struggle to judge body orientation and movement limits while scoring lower on counterfactual than on descriptive questions. To understand how to improve human-centric spatio-physics reasoning, we compare three representations (Image, Graph, and Code) and two states, estimated from the image or taken from the rendered scene. We learn two best lessons for human-centric spatio-physics reasoning: (1) Structured body information can improve spatial judgments, but physical reasoning also needs reliable physical state relevant to the question. (2) Computing the judgment from that state can improve physical reasoning when they use reliable scene geometry and physical properties. SPARKBench with the best lessons provides a basis for identifying model weaknesses and designing better representations and computations for human-centric spatio-physics reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.