Struct2Real: A Systematic Framework for Accurate and Efficient Structure-Grounded Object Image Generation
Abstract
Recent advances in image generation have enabled the creation of high-quality visual content with impressive semantic fidelity. However, generating object images under fine-grained structural constraints, particularly preserving topology and spatial layout, remains an open challenge. We propose Struct2Real, a structure-grounded generation framework that enables photorealistic object image synthesis under explicit structural control while supporting flexible viewpoint and structural editing, consisting of twofold. 1) we develop a structure modeling system that enables users to create a 3D structural representation named StructMap, which is composed of various geometric primitives and allows precise, intuitive and editable structural specification. 2) We develop the Structure-Aware Generation Agent (SAGA), a skill-oriented agent that formulates structure-grounded image synthesis as a dynamic skill orchestration process and enables photorealistic object image generation under structural constraints encoded in StructMap. SAGA further incorporates a realism memory module that retains visual attributes throughout generation, enabling viewpoint and structural editing with consistent visual attributes. Extensive experiments demonstrate that Struct2Real achieves strong performance in structure-grounded object image generation while ensuring low user effort required for this task, highlighting the practicality and effectiveness of our method. Please refer to more details in the Appendix and supplementary material.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.