acceptodds
Under review as a conference paper at ICLR 2027

AssemblyVerse: Towards Multimodal 3D Assembly Reasoning & Planning

Abstract

The task of 3D assembly is not merely predicting a sequence of independent part poses; it is a long-horizon, state-conditioned decision problem that couples reasoning about what to assemble next with planning how to move each part into place. Existing benchmarks address either 6-DoF pose reasoning or trajectory prediction in specialized regimes, leaving open the general process-level learning across domains and initial configurations. Towards this end, we introduce AssemblyVerse, a large-scale, multi-domain dataset for process-level multimodal assembly reasoning and planning of diverse objects and complex assembly configurations. AssemblyVerse contains 10,809 assemblies spanning five domains and 192,412 assembly steps. Each step is aligned with rendered instruction diagrams, text instructions, evolving assembly states, and geometrically verified collision-free trajectories. Building on AssemblyVerse, we present AssemblyMind, a vision-language 3D assembly model that grounds each subsequent assembly step through an auxiliary order-grounding head and autoregressively generates 6-DoF trajectories as reordered discrete pose tokens. Our staging-randomized supervised fine-tuning improves robustness to diverse initial part configurations, while GV-GRPO directly optimizes geometric validity of predicted trajectories without sacrificing final-pose accuracy. Extensive experiments across domains demonstrate that strong performance on AssemblyVerse enables transferability of process-level learning. Compared to the previous state of the art, AssemblyMind substantially improves order grounding, final-pose prediction, and collision-free trajectory generation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.