CognPhys: A Benchmark for Inductive Physical Reasoning in Vision-Language Models
Abstract
Existing benchmarks for physical reasoning in Vision-Language Models (VLMs) primarily assess rule following—predicting outcomes governed by known Newtonian laws—yet a hallmark of human physical cognition is rule discovery: observing unfamiliar phenomena and inducing the underlying principles from scratch. This capability is essential for embodied agents adapting to novel environments, robot control, and automated scientific discovery. We introduce CognPhys, a diagnostic benchmark comprising multiple physics worlds, each governed by distinct rules that systematically deviates from real-world physics. We evaluate a broad suite of open-source and closed-source VLMs using a structured scoring protocol. Our extensive evaluations reveal a profound deficiency in state-of-the-art VLMs: they are severely bottlenecked at rule induction. Even the most capable models consistently fail to synthesize coherent rules from observed motion, constrained by an over-reliance on memorized physical priors. CognPhys uncovers a fundamental blind spot in current VLM physical reasoning and establishes a challenging testbed for scientific and embodied AI.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.