VisCAD: A foundation model suite with multimodal industrial CAD intelligence
Abstract
AI-assisted computer-aided design (CAD) for industrial products involves two challenging phases. Part-level generation maps diverse forms of user intent, including renders, text descriptions, 2D drawings, and real photographs, to executable programs in a CAD domain-specific language. Assembly-level generation must additionally handle interacting parts, plan mating relations, estimate poses, and place all parts correctly. Existing specialized CAD models are commonly trained on narrow input domains (renders or texts) and often generalize poorly, while general-purpose frontier models cover broader inputs but perform inconsistently across CAD domains. We present \viscad, a foundation model suite designed to provide both broad generalization and strong CAD capability for realistic industrial products. In its core is \viscadm, a 27B model trained through mid-training and post-training for part-level design generation. On \pcb and \rcb, \viscadm achieves the highest average part-level score among the evaluated models, reaching 0.5540 compared with 0.5496 for the strongest frontier model. Reusing \viscadm as a test-time verifier can further raise the score to 0.5797, an approximately 5 percent relative improvement over the previous state of the art. \viscad also includes a domain-specific harness that leverages frontier models for complex assembly generation and demonstrates advantages over general-purpose harnesses in both quantitative and qualitative evaluations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.