acceptodds
Under review as a conference paper at ICLR 2027

World As a Program: Programmable Spatial Reasoning for Multimodal LLMs

Abstract

Although MLLMs have advanced rapidly, spatial reasoning remains limited by implicit, geometrically unconstrained representations, while explicit tool-based approaches often rely on external perception modules for object grounding and geometric acquisition, underusing native visual-semantic capabilities. Recent advances in 3D reconstruction can recover explicit 3D representations of real-world environments, where scene entities are directly accessible and manipulable, suggesting 3D assets as a natural programmable substrate for spatial reasoning. This motivates a natural question: Can explicit 3D world states be made directly programmable for MLLMs? To this end, we propose WoRld As a Program (WRAP), a framework that converts standard video observations into an addressable and executable 3D scene state for MLLMs, where visual objects are exposed as Grounded Variables for 3D Program Reasoning. WRAP explicitly binds visual object instances to their 3D entities as grounded variables, then enables the MLLM to generate executable reasoning programs over these variables while delegating only deterministic computation to an external executor. Extensive experiments on both video spatial reasoning and 3D scene editing benchmarks demonstrate the effectiveness and generality of WRAP. Our method achieves state-of-the-art (SOTA) performance on ReVSI, improving video spatial reasoning accuracy by up to 5.2% over prior SOTA methods, while also outperforming existing approaches on EditRoom-DB test set, achieving an absolute IoU improvement of 0.19 over prior methods. These results demonstrate the programmable 3D representations benefit both video spatial reasoning and scene manipulation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.