acceptodds
Under review as a conference paper at ICLR 2027

ShapeMind: Native 3D Understanding and Generation of Semantically Structured 3D Assets

Abstract

Production-ready 3D assets demand more than visual fidelity: they require semantically structured geometry for editing, composition, and interaction in applications such as embodied simulation. We present ShapeMind, a framework that generates low-poly 3D assets with explicit semantic part structure and, where applicable, kinematic structure. Following the principle of “think before you generate”, ShapeMind couples ShapeVLM for 3D semantic understanding with Semantic PartGen for native mesh generation. ShapeVLM combines visual evidence with coarse geometry encoded as compact text within the existing vocabulary of a vision-language model (VLM). This design reuses pretrained semantic and commonsense knowledge for 3D structural reasoning without introducing a heterogeneous modality. The resulting understanding of object composition and kinematics provides the structural basis for Semantic PartGen to synthesize high-quality low-poly meshes organized into meaningful parts. Its part-aware vertex and topology generation jointly models the object while preserving each part as an independently editable mesh. To support this conditional generation task, we curate 300K semantically annotated low-poly assets. Experiments demonstrate state-of-the-art results in both 3D object semantic understanding and structurally coherent part geometry generation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.