WBench: A Comprehensive Multi-turn Benchmark for Interactive World Model Evaluation
Abstract
Interactive world models are advancing rapidly, yet existing benchmarks cover only part of the required competencies, leaving no unified standard for systematic evaluation. To fill this gap, we introduce **WBench**, a comprehensive multi-turn benchmark for interactive world model evaluation along five dimensions, namely video quality, setting adherence, interaction adherence, consistency, and physics compliance. WBench contains 289 test cases and 1,058 interaction turns, where each case specifies a world setting and a multi-turn interaction sequence, covering diverse scenes, styles, subjects, and both first- and third-person perspectives, together with four interaction types, including navigation, subject action, event editing, and perspective switching. For navigation, WBench unifies text, 6-DoF pose, and discrete-action control, enabling evaluation of models with different native input interfaces. Evaluation uses 22 automatic sub-metrics that combine specialist vision models with vision-language models. Across 20 state-of-the-art models, we find that no single model performs strongly across all dimensions. We hope these diagnostic insights into model strengths, weaknesses, and open challenges will help advance research on interactive world models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.