3Dify-Anything: A Unified Model for Multimodal 3D Generation via Context Pre-Training
Abstract
We introduce 3Dify-Anything, a single model that turns diverse inputs into high-quality 3D assets with fast inference. Inputs can be text, images, sketches, depth maps, point clouds, voxels, bounding boxes, or any combination of these. The model rests on two ideas. The first is an encoder-decoder design that separates understanding the context from generating the shape. A large transformer context encoder reads the input once, and a lightweight rectified-flow decoder generates the 3D latent in a few ODE steps while attending to the encoder's representation. Because only the decoder runs at every step, capacity is concentrated where it is cheap, giving fast inference without sacrificing fidelity. The second idea is context pretraining with cross-modal masking, which makes the encoder strong enough for a small decoder to rely on. The encoder is pretrained with masked token prediction on single- and multi-modal contexts, including the target 3D shape tokens. For multi-modal contexts, one modality is kept visible while the other is heavily masked, so the encoder learns to predict one modality from another. Across a wide range of settings, 3Dify-Anything is competitive with specialized single-task baselines while offering far greater flexibility in how inputs are composed. A compact variant enables near real-time multimodal 3D generation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.