acceptodds
Under review as a conference paper at ICLR 2027

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Abstract

A 3D scene reconstructed from a single image is most useful when represented not as a rendering or a fixed 3D output, but as an explicit scene program whose execution yields a scene that can be inspected, edited, and queried. We present **LEGO-Anything**, an *Image-to-Code* framework in which a coding agent builds such a program by iteratively writing and executing Blender code, inspecting the evolving scene and its renderings, and revising the program. To evaluate how well such agents recover scenes end to end, we introduce **LEGO-Bench**, a simulator-grounded benchmark with 208 images from 104 diverse scenes spanning indoor and outdoor environments. LEGO-Bench separately scores artifact validity, visible-surface geometry, and rendered appearance, and its simulator-grounded design makes the benchmark naturally extensible while retaining precise automatic evaluation. Across the evaluated coding agents, GPT-6-astra achieves the strongest overall results, with 53.4% indoor and 39.6% outdoor scores, yet substantial gaps remain between delivering valid scene artifacts and faithfully recovering scene geometry and appearance. To understand these failures, we analyze agent construction trajectories and identify three recurring issues: weak scene initialization, regressive edits during iteration, and unreliable self-evaluation. These findings motivate **LEGO-Plugin**, a training-free harness plugin for more controlled iterative scene construction, which improves all six evaluated models, with relative gains of up to 62.7% in overall score. Finally, we ask whether reconstructed scenes are faithful enough to represent natural images and support vision tasks. In **LEGO-World**, we take scenes reconstructed by GPT-6-astra and derive object detections, instance masks, and relative depth as deterministic queries on each scene. These readouts show non-trivial performance across all three tasks but fall well short of specialized vision models, suggesting that program-constructed scenes from current coding agents are a promising but not yet sufficiently precise representation of natural images.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.