acceptodds
Under review as a conference paper at ICLR 2027

What-If World: A Contrastive Reasoning Benchmark for Video World Models in Embodied Scenarios

Abstract

Video generation models are increasingly used as world simulators for driving and robotic manipulation, where their value lies in answering what-if questions: if the car brakes harder, or the object is heavier, what happens next? Answering such questions requires a model to change its output correctly when its input changes. Many existing benchmarks score each video in isolation rather than test this contrast: a model can generate two videos that each look plausible yet are wrong relative to each other. We introduce What-If World, a contrastive reasoning benchmark of 319 prompt pairs grounded in real frames from nuScenes and DROID. The two prompts in each pair describe the same scene with a single physical variable altered, drawn from a taxonomy of six variables shared across driving and manipulation. The prompts differ only slightly in wording, but the correct videos should diverge substantially. We score each pair with APEO, a rubric with four checks: whether each video follows its prompt (Adherence), whether its motion is physically consistent (Physics), whether the shared scene is preserved (Environment), and whether the two videos differ in the expected outcome (Outcome). Across nine state-of-the-art models, contrastive reasoning remains challenging: none exceeds 53% on the paired score, and the four open-source models average 28.6%. This failure is easy to miss when each video looks plausible on its own. Every model scores 28 to 49 points lower on paired physics than on single-video physics, a gap we call the contrastive bottleneck: a common failure is to produce similar trajectories under both prompts. Beyond being hard to see, the failure falls on the what-if questions that matter for embodied agents; for example, models are often far better at making things move than at making them stop. Realism therefore does not imply contrastive reasoning, which What-If World measures one intervention at a time, a useful test before a video generation model serves as a world simulator.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.