acceptodds
Under review as a conference paper at ICLR 2027

ISArt: Single-Image, Structure-Guided Generation of Sim-Ready Articulated Assets

Abstract

Generating simulation-ready articulated assets from a single image requires detailed part geometry and coherent kinematic structures. While Coding agents built on models such as GPT6-Astra can automate programmatic modeling and joint specification, faithfully reproducing intricate, image-specific part geometry through code remains challenging. We present ISArt, a closed-loop framework that combines agentic structural reasoning with learned part-level geometry generation. Given an input image, a vision-language model (VLM) iteratively decomposes the object into an editable structural representation, identifying its semantic parts, spatial extents, joint types, and parent–child relationships. For objects with concealed interiors, ISArt additionally synthesizes open-state references to reveal these structures. Conditioned on both visual observations and the inferred structure, a diffusion transformer generates detailed geometry for each part, followed by part-level texture synthesis. An articulation agent then estimates joint origins, axes, and motion ranges from the generated geometry and evaluates geometric and motion consistency. These feedbacks guide iterative refinement, fallback selection, and candidate acceptance or rejection. By assigning detailed geometry synthesis to a learned generative model and structural planning and articulation refinement to an agentic workflow, ISArt produces textured part meshes with URDF-defined kinematics for robotic simulation and interactive applications.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.