PROWBench: An Extensible Benchmark for Programmable World Models
Abstract
Programmable world models separate executable dynamics from visual generation, offering a promising foundation for next-generation game engines. However, their visual adherence to explicit rules and interactions remains insufficiently evaluated. Existing benchmarks emphasize video quality and camera controllability, with limited support for evaluation grounded in replayable world records. We introduce PROWBench, comprising 170 programmatically constructed episodes and 600 proxy videos covering diverse scenes and interactions. An extensible framework constructs scenes, controls behaviors, and renders three synchronized representations per camera view, including coarse 3D, semantic proxy, and colored oriented bounding boxes, from shared state and interaction records. The benchmark covers first- and third-person perspectives, with synchronized multi-view observations available for a subset of episodes. Two VLM-based metrics, Logic–Render Alignment and Interaction Success Rate, assess prompt-rule adherence and the visual realization of timestamped engine-recorded events, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.