FaceFlow: Native Artist-like Mesh Generation with Face Flow Diffusion
Abstract
Diffusion-based artist-like mesh generation has recently emerged as an important direction in 3D content creation. However, the discrete nature of mesh topology poses a fundamental challenge for continuous diffusion models. In this paper, we further investigate the parallel generation of complete discrete mesh sequences. To this end, we propose FaceFlow, a diffusion-based framework for artist-like mesh generation. FaceFlow employs a Face VAE to encode mesh faces into compact face latents and decode them into complete mesh sequences in a single step. However, conventional sorting strategies cannot resolve the assignment ambiguity caused by the variable cardinality and non-uniform spatial distribution of face latents. Consequently, the face diffusion model cannot reliably determine which target face latent each noisy latent should represent. To address this issue, we introduce a Point Sink Flow model that transports a dense noisy point cloud into a sparse set of face centers. These centers specify the 3D locations of the corresponding face latents and are injected into the Face DiT as positional embeddings, explicitly establishing spatial correspondences between noisy and target latents. This spatial guidance enables the Face DiT to learn a more accurate transformation toward the face-latent distribution, thereby improving the quality of the generated mesh sequences. Experiments on Toys4K demonstrate that FaceFlow achieves state-of-the-art performance in both mesh reconstruction and generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.