IMSCENE: AN AGENTIC FRAMEWORK FOR RECONSTRUCTING SIMULATION-READY 3D SCENES FROM A SINGLE IMAGE
Abstract
Physically reliable 3D scenes are essential for embodied navigation, robotic manipulation, and interactive environment simulation. Existing image-to-3D methods generate plausible objects and recover scene layouts, but inconsistencies between generated geometry, estimated poses, and inter-object relationships can still lead to interpenetration, floating objects, and unstable support. Resolving these inconsistencies from a single image requires reconciling incomplete visual evidence with geometric and physical constraints on the reconstructed scene. We present ImScene, an agentic framework for reconstructing editable 3D scenes from a single image. A detection agent corrects missing, duplicate, and mislabeled instances, while a vision-language model constructs a physical scene graph encoding support, containment, and attachment relations. Next, we reconstruct a scene point cloud from estimated depth and establish a gravity-aligned metric coordinate system. Guided by the scene graph, a differentiable optimizer then jointly refines the poses and scales of objects generated by an image-to-3D model, using image and depth alignment, surface contact constraints, and penalties for interpenetration and unsupported placement. Finally, a simulator-guided agent identifies unstable object placements from simulated motion and support contacts, then applies bounded pose adjustments or rigid-body settling to improve stability while maintaining consistency with the input image. Comprehensive comparisons demonstrate improved consistency with the input image and greater physical plausibility, providing a stronger foundation for subsequent scene editing and simulation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.