acceptodds
Under review as a conference paper at ICLR 2027

PhySimCode: A Benchmark and Evaluation Method for Physics Video to Code Generation

Abstract

Multimodal large language models (MLLMs) are good at video understanding and code generation, but it remains unclear whether they can recover the physical laws underlying videos, a capability essential for converting videos into executable simulations and for providing structured physical supervision in robotics. Captioning motion video does not imply that a model has inferred its governing dynamics, estimated quantitative parameters, or can regenerate the observed scene from code. Existing video–physics benchmarks evaluate plausibility, property estimation, or future-state prediction, but offer limited insight into whether models can translate visual dynamics into runnable simulation code. We introduce **PhySimCode**, a benchmark for physics video-to-code generation: given only a raw video, a model must infer the relevant physics, recover parameter values, and generate runnable Python that reproduces the input video. The dataset contains **160.6K procedurally generated samples** spanning **162 physics phenomena** across analytic SciPy simulations and PyBullet-rendered rigid-body, articulated, fluid, and contact-rich scenes; each sample bundles the source simulation code, sampled physical parameters, rendered video, and a structured chain-of-thought covering observations, laws, ODEs, numerical methods, and validation checks. We evaluate ten state-of-the-art closed- and open-source MLLMs along physical-law correctness, simulation equivalence, code executability, and parameter recovery. Models reliably state physics laws (4.27–4.93/5) yet score only 1.51–2.47/5 on inferring them from video, and the best system (Claude Sonnet 4.6) recovers just 7.7% of ground-truth parameters within , exposing a substantial video-to-simulation gap and positioning PhySimCode as a testbed for multimodal systems that translate perception into precise, executable, and scientifically verifiable physics. [Dataset](https://kaggle.com/datasets/f38d602eae34cadbac5ed3585739c8e92c8897a040d50d18c0992e96659d53d2) and [Code](https://github.com/PhysicsBenchmark/physics-video-to-code.git) for generation and evaluation are released.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.