TOPOS: High-Fidelity and Efficient Industry-Grade 3D Head Generation
Abstract
Industrial 3D head generation requires fixed reference topology for production-level rigging, skinning and animation, yet existing 3D generative models typically produce inconsistent mesh topologies which are incompatible with production pipelines. In this paper, we present TOPOS, a single image conditioned framework that jointly generates high-fidelity geometry and appearance under a unified industry-standard topology. To model heads under this unified topology, we proposed a novel variational autoencoder structure, termed TOPOS-VAE. Inspired by multi-modal large language models (MLLMs), TOPOS-VAE leverages the Perceiver Resampler to convert input pointclouds sampled from head meshes of diverse topologies into the target reference topology. Building upon TOPOS-VAE's structured latent space, we train a rectified flow transformer, TOPOS-DiT, to efficiently generate high-fidelity head meshes from a single image. We further present TOPOS-Texture, an end-to-end module that produces relightable and geometry-aligned UV texture maps from the same portrait image via fine-tuning a multi-modal image generative model. Extensive experiments demonstrate that TOPOS achieves state-of-the-art performance on 3D head generation, surpassing both classical face reconstruction methods and general 3D object generative models, highlighting its effectiveness for digital human creation. We will release our code and trained models to facilitate future research.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.