acceptodds
Under review as a conference paper at ICLR 2027

PhysCot: A Benchmark for Multi-Image Physical Reasoning

Abstract

Physical reasoning requires models to track ordered states, infer dynamics, and predict physically valid futures. We introduce PhysCot, a benchmark of 400 physics-grounded visual sequence-completion problems for multimodal large language models. Each problem provides four simulation frames from one of eight subfields and asks the model to select the fifth from four visually similar candidates generated by physical-state perturbations. PhysCot combines coherent sequences, image candidates, traceable states, and controlled distractors. Because candidates differ in rendered physical state rather than answer text, the task is designed to evaluate state extraction and consistency checking under a shared explicit physics scaffold. Its three protocols evaluate full visual reasoning, focus on visual state perception with reduced temporal context (PhysCot), and probe physical modeling with textual candidate states (PhysCot). These protocols are diagnostic rather than causal interventions and provide a common basis for comparing model failure patterns across physical domains. Within the fixed synthetic simulator–renderer distribution, no model exceeds 50% on the full protocol, while the strongest model reaches 78.56% in the mixed text-augmented condition. Temporal-context effects are heterogeneous across models and input configurations; these results characterize in-distribution behavior and do not establish robustness to held-out parameters, renderings, or real imagery outside the benchmark's controlled evaluation setting.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.