acceptodds
Under review as a conference paper at ICLR 2027

Phys-WorldBench: A Unified Benchmark for Evaluating Embodied World Models from Passive Prediction to Active Intervention

Abstract

Embodied intelligence requires model to understand how actions transform the physical world. World models provide a natural foundation for this ability, yet their physical interaction capabilities remain insufficiently characterized. Existing benchmarks often evaluate passive prediction or action-conditioned modeling separately, lacking a unified protocol that connects world prediction, action-effect modeling, and real-world action generation. They also commonly organize evaluation around semantic tasks, making it difficult to cover the underlying mechanisms of physical interaction. We introduce , a unified real-world benchmark that organizes tabletop manipulation through interaction skills and physical constraints. Its three-level protocol spans passive future prediction, action-conditioned world modeling, and world action modeling within a shared physical interaction space. Built on 4,500 real-robot trajectories (30 hours), Phys-WorldBench covers 50 atomic, compositional, and out-of-distribution tasks. Across representative world models and embodied policies, we observe substantial variation across the three levels, with persistent challenges on compositional and out-of-distribution tasks. Visually plausible prediction does not consistently translate into accurate action-effect modeling or robust real-world execution. Cross-level analysis further finds significant moderate associations between adjacent levels (– for L1–L2 and – for L2–L3, all across Wan and Cosmos), while passive prediction remains weakly associated with real-robot progress. Together, these results suggest that predictive quality alone provides an incomplete measure of embodied world-model capability, motivating evaluation across the full progression from world prediction to physical interaction.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.