The Next Step in 3D Generation Post-Training
Abstract
Post-training has become an effective approach for aligning generative models with human preferences and task-specific objectives, yet its application to 3D generation remains limited. Existing methods primarily target text-to-3D generation, support only specific model architectures, and rely on insufficiently justified reward and rendering designs, often yielding marginal improvements at substantial training cost. To address these limitations, we introduce Next3DGen, an unified framework that takes the next step in 3D generation post-training by supporting both text-conditioned and image-conditioned generation post-training across single-stage and multi-stage models. We systematically investigate recent efficient RL algorithms and develop a principled reward pipeline based on explicit criteria for 3D generation quality. For image-to-3D generation, Next3DGen applies complementary rewards to reference-aligned and its antipodal view, improving unobserved regions while preserving input alignment. For multi-stage generators, we propose tree-structured rollouts that jointly optimize geometry and texture through stage-specific credit assignment. Experiments on text-to-3D and image-to-3D benchmarks show that Next3DGen consistently improves generation quality, condition alignment, and 3D consistency over pretrained models, substantially outperforms existing post-training methods, while maintaining training efficiency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.