PhysCot: A Benchmark for Multi-Image Physical Reasoning
Abstract
Physical reasoning requires models to track ordered states, infer dynamics, and predict physically valid futures. We introduce PhysCot, a benchmark of 400 physics-grounded visual sequence-completion problems for multimodal large language models. Each problem provides four simulation frames from one of eight subfields and asks the model to select the fifth from four visually similar candidates generated by physical-state perturbations. PhysCot combines coherent sequences, image candidates, traceable states, and controlled distractors. Because candidates differ in rendered physical state rather than answer text, the task is designed to evaluate state extraction and consistency checking under a shared explicit physics scaffold. Its three protocols evaluate full visual reasoning, focus on visual state perception with reduced temporal context (PhysCot), and probe physical modeling with textual candidate states (PhysCot). These protocols are diagnostic rather than causal interventions and provide a common basis for comparing model failure patterns across physical domains. Within the fixed synthetic simulator–renderer distribution, no model exceeds 50% on the full protocol, while the strongest model reaches 78.56% in the mixed text-augmented condition. Temporal-context effects are heterogeneous across models and input configurations; these results characterize in-distribution behavior and do not establish robustness to held-out parameters, renderings, or real imagery outside the benchmark's controlled evaluation setting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.