acceptodds
Under review as a conference paper at ICLR 2027

ExeWorldBench: Benchmarking Code-Based Executable 3D World Modeling

Abstract

Coding agents are increasingly evolving from general-purpose programming tools into a means for constructing and simulating executable worlds. Existing approaches have begun to explore this direction from two perspectives. The first takes explicit textual descriptions as input and turns them into executable world that exhibit the specified behaviors. The second begins with visual observations and seeks to infer the hidden physical properties, dynamics, and structural organization that explain the world. However, executable world modeling goes beyond either paradigm alone, requiring both world understanding and transferring that understanding into executable code-based construction. Yet, this integrated capability remains insufficiently evaluated. We introduce ExeWorldBench, the first benchmark for evaluating this capability across three complementary tasks.Mechanism Executability evaluates executable realization of individual mechanisms from visual observations. Compositional Executability extends this to multi-stage systems with spatial and causal dependencies. Intent-Driven Executable World Synthesis requires models to construct executable worlds from high-level goals and rules. Together, these tasks form a progressive evaluation hierarchy with increasing complexity and autonomy. Experiments on frontier models reveal substantial gaps in executable world modeling, underscoring the challenge of integrating world understanding with executable world construction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.