acceptodds
Under review as a conference paper at ICLR 2027

FLiTo: Fast High-Fidelity 3D Generation via Image Supervision

Abstract

Recent advances in latent 3D generative modeling have enabled high-quality 3D content generation from images. However, state-of-the-art models typically require many generation steps, resulting in substantial inference cost and limiting their use in interactive applications. While distillation can reduce the number of sampling steps, generation quality degrades significantly in the few-step regime, particularly for single-step generation. In this work, we study how 2D supervision can be effectively incorporated to improve single- and few-step latent 3D generation. We derive a variational objective that naturally unifies conditioning-image reconstruction with distribution matching against a teacher 3D generator. Building on this formulation, we incorporate complementary forms of 2D supervision, including feature-distribution matching, semantic consistency, and human preference. Through extensive ablations, we analyze the contribution of each supervision signal and show that 2D supervision is particularly effective in the challenging single-step regime. Our final single-step generator achieves state-of-the-art speed and generation quality among single-step 3D generators, reducing FID by 47% over the strongest baseline while running in 67.5 ms on an NVIDIA B200 GPU and 2.4 seconds on an MacBook Pro with M5 Max.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.