Phi-Art: Physics-Aware Articulated Object Generation from a Single Image with a Multimodal Large Language Model
Abstract
Generating articulated assets from a single image requires more than reconstructing visible surfaces. Part geometry, kinematic structure, and physical attributes must be inferred from incomplete visual evidence to support interactive simulation. How- ever, existing models are struggling to tackle these problem all together. Thus, we introduce Phi-Art: -Art, a multimodal framework that couples structured object prediction with parallel geometry decoding. A multimodal language model predicts part semantics, 3D bounding boxes, joint parameters, and physical attributes, providing a coarse spatial plan for asset generation. A Qformer-style decoder combines the model’s hidden representations with image features to predict whole-object geometry codes in parallel, separating geometric decoding from autoregressive text generation. We further corrupt structured descriptions during training to encourage the use of visual evidence and reduce reliance on accurate textual layouts. The predicted geometry and layouts guide part synthesis, followed by kinematic assembly and export to simulator formats. Experiments on PhysXNet and Mobility show that our method outperforms the existing baselines in the geometry accuracy, training and inference efficiency with a competitive physical attributes prediction. We will release our code, model checkpoints, and processed training data
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.