VoxProg: 3D Spatial Reasoning through Voxel Programming
Abstract
Language models' abilities to interpret tasks and generate programs offer a promising basis for 3D spatial reasoning. Yet applying these abilities to spatial instructions requires generated programs to construct geometric references that language alone does not fully specify. In this work, we introduce \method, which addresses this challenge through voxel programming: expressing task interpretations as executable programs that operate on local voxel views of scene surfaces. By organizing surface geometry on regular grids with explicit coordinates and neighborhood relations, these programs construct task-relevant regions and boundaries while retaining spatial extent and shape. A downstream pose program then uses this geometry to compute target object poses, given explicit scene geometry and known initial object poses. Visual and numerical execution feedback allows the model to inspect intermediate structures and iteratively revise its programs. We also introduce \suite, a benchmark of 71 open-ended placement tasks whose diverse and complex spatial requirements must be inferred from the interplay of instructions, scene contexts, and object geometries. Its evaluation combines Intent Alignment, measured by image comparison with the ground-truth placement, with Physical Validity checks. Code will be released upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.