TopCode: Disentangling Visual Evidence and Physical Dependencies for Executable 3D Scene Synthesis
Abstract
Generating executable 3D indoor scenes from top-down references supports interior design and embodied intelligence. However, existing methods struggle to reconcile visual fidelity with physical completeness while preserving unrelated content during revision. We propose TopCode, an agent harness that translates Top-down images into executable scene Code by disentangling physical dependencies from visual evidence. Scene Memory preserves accepted attributes and their evidence in DeltaTree, while a separate Coupling Graph captures physical dependencies. Incremental scene compilation in Blender completes unobserved geometry and articulation, accepting updates only after visual and execution checks pass. When checks fail, counterexample-guided revision uses dependencies and hierarchical pose composition to localize changes while preserving accepted values with valid visual support. Across 100 top-down references and 120 local editing tasks, TopCode improves reference consistency, scene validity, and preservation of unrelated content over evaluated baselines. Generated scenes also improve navigation through post-training and support simulated manipulation, demonstrating their utility for embodied tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.