MUSE: Benchmarking Assembly Structure and Interface Geometry in Multi-Component Text-to-CAD Generation
Abstract
Large language models (LLMs) have recently advanced text-driven 3D generation, yet generating CAD models that satisfy practical design requirements remains challenging. Existing benchmarks focus primarily on single-part CAD models and geometric similarity, providing limited assessment of assembly structure and interface geometry. To address this gap, we introduce \model, a Text-to-CAD benchmark for generating multi-component, editable boundary representation (B-Rep) assemblies. \model pairs practical design instances with structured Design Specifications and evaluates generated outputs across four dimensions: code executability, geometric validity, structural and interface fidelity, and functional consistency. We further introduce Structure and Interface Evaluation (SIE), which combines reference Physical Assembly Graphs with direct B-Rep analysis to assess component correspondence, connectivity, interface types, and fit clearances. Experiments on proprietary and open-weight LLMs show that high execution success does not necessarily translate into correct assembly structure or interface geometry. Together, \model provides a benchmark and evaluation framework for studying multi-component Text-to-CAD generation beyond shape similarity. Our dataset and code are available at https://anonymous.4open.science/status/muse_-8C3B.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.