acceptodds
Under review as a conference paper at ICLR 2027

PartGround: Benchmarking, Analyzing, and Improving VLM-Based 3D Generation

Abstract

Vision-Language Models (VLMs) can generate editable 3D assets through code, complementing diffusion-based methods. Whole-object similarity alone cannot establish component, material, and relational fidelity in an assembled asset. To address this evaluation need, we introduce PartGround-Bench for reconstruction from reference images and supplied part specifications, with 1,195 manually annotated objects selected from 82,717 Blendkit models across 12 category groups. Fixed part correspondences and shared whole-object alignment enable direct scoring of overall shape, part geometry and completeness, per-part materials, and spatial relationships, preserving placement errors. We evaluate nine single-call configurations on the full benchmark and five agent configurations on a subset. Under Visual Part Guidance, two agents achieve similar mean IoU (0.370 and 0.372) but different mean Part-F1 (0.149 and 0.168). Higher reasoning effort accompanies better single-call geometry scores, without uniformly improving materials or reducing inter-part penetration. We further evaluate a reusable multi-agent modeling skill with shared planning and independent review, improving overall and part geometry, layout, and contact consistency across three agent configurations. For Kimi-K3 (max) under Visual Part Guidance, our relative gains are 5.9% in whole-object IoU, 17.4% in Part-F1, 3.2% in layout \(\DeltaAUC\), and 8.2% in ContactF-AUC, demonstrating the effectiveness of our multi-agent modeling skill in improving geometric and relational fidelity.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.