acceptodds
Under review as a conference paper at ICLR 2027

Does One Prompt Represent a Physical Scenario? Physics Evaluation through Prompt Sets

Abstract

Evaluating whether video generation models understand the physical world is essential, since physical understanding is a prerequisite for their use as world models. Existing physics benchmarks generate a video from a text prompt describing a physical scenario and score its physical plausibility. However, each scenario is represented by a single prompt, although it can be described in many semantically and physically equivalent ways. We show that this single-prompt design makes the evaluation sensitive to the choice of prompt. Using a large language model, we rewrite benchmark prompts while preserving their semantics and physics, and evaluate five video generation models on VideoPhy-2 and PhyGenBench with fixed seeds. Rewriting alone changes the physical score in 35.5% and 35.7% of cases, flips the official pass decision in 13.3%, and changes the point ranking of 4 of the 5 models on each benchmark. To address this prompt sensitivity, we propose Prompt Set Evaluation, which scores each scenario over a set of validated equivalent prompts and reports the mean score with its variation; five prompts per scenario substantially improve the stability of scores and rankings. The sensitivity is not random noise, since it persists under fully repeatable evaluators. In a blinded human study of video pairs whose scores differ, raters seldom share the evaluator's preference and judge about half of the pairs equally plausible, which suggests that the vision-language models serving as evaluators also respond to differences between videos unrelated to physics. Physical understanding should therefore be evaluated over a distribution of descriptions rather than a single prompt.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.