acceptodds
Under review as a conference paper at ICLR 2027

CRAFT: ContRollable Generation of Articulated Objects From Image and Text

Abstract

Controllable articulated-object generation is important for building interactive 3D worlds, supporting embodied intelligence and robotic simulation, and creating functional digital assets. A reference image reveals what an object looks like, but not necessarily how a user wants its parts to move. We introduce CRAFT, a framework for visually grounded, controllable articulated-object generation from a reference image and a textual articulation specification. CRAFT separates coarse structural control from detailed asset synthesis through an editable articulated 3D proxy that represents part dimensions and positions, spatial relationships, and kinematic parameters. We first construct an initial proxy using a category-informed structural initialization strategy. We then refine its spatial arrangement through adaptive structural optimization that combines instance-conditioned geometric objectives with semantic guidance from an image-derived 3D estimate. Finally, we generate detailed part assets and assemble them according to the refined proxy. Experiments demonstrate that CRAFT generates articulated objects consistent with the reference appearance while enabling text-specified control over part structure and motion.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.