Swift3D: End-to-End Acceleration of High-Resolution Native 3D Generation
Abstract
Native 3D generative models aim to synthesize detailed geometry and textures from a single image, but producing high-resolution textured meshes remains time-consuming, particularly during iterative diffusion sampling, geometry decoding, and UV parameterization. To address these bottlenecks, we present Swift3D, an end-to-end acceleration framework for high-resolution native 3D generation. We develop stage-aware distillation of diffusion transformers (DiTs), combining phased consistency model (PCM) initialization with teacher anchoring of coarse geometry during distributional and adversarial refinement, enabling generation with fewer sampling steps. Adaptive geometry decoding uses probes to guide subdivision and reconciles coarse and fine fields through consistent reconstruction, reducing redundant neural field queries for high-resolution surface extraction. Progressive UV atlasing aligns chart construction with projection directions and iteratively refines only failed charts, reducing UV parameterization overhead on dense meshes. Experiments on geometry reconstruction, geometry generation, texturing, and end-to-end generation demonstrate substantial acceleration while retaining high fidelity. Swift3D reduces average end-to-end latency from 520.56 to 47.39 seconds (), outperforming the evaluated open-source baselines in asset quality. Further distilling the geometry-refinement and texturing DiTs to one step each enables generation in under 30 seconds with quality comparable to high-quality open-source state-of-the-art methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.