ArtiCritic: A Vision–Language Critic for Generated Articulated Objects
Abstract
Despite recent advances in articulated object generation, systematic evaluation of generated assets remains underexploredRecent breakthroughs in vision-language models (VLMs) have made it tempting to simply ask a VLM to judge an asset holistically, but such judgments tend to fall short on two fronts: they struggle to localize problems precisely in the image, and they reason poorly about fine-grained structure. We introduce ArtiCritic, a training-free, VLM-assisted framework that combines three-stage inspection with an overall review. More specifically, ArtiCritic decomposes evaluation into evidence-guided checks that progress from intrinsic validity to reference-based assessment of part completeness and articulation plausibility. Because reference objects and generated assets may represent structure at different levels of granularity, we introduce cross-granularity structural matching (CGSM), which associates reference parts with sets of candidate surface regions and checks their expected relative motions against corresponding rigid groups and joint paths. Each inspection localizes potential defects to specific components and records the supporting visual and structural evidence. A final VLM adjudicator then reviews these findings, filters unsupported claims, consolidates confirmed defects, and produces an overall assessment together with targeted repair suggestions. Experiments on manually constructed test cases show that ArtiCritic can reliably identify defective articulated assets while providing localized and actionable diagnostic feedback.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.