acceptodds
Under review as a conference paper at ICLR 2027

Hybrid3D-RL: Hybrid-Space Reinforcement Learning for High-Fidelity 3D Generation

Abstract

Reinforcement-learning post-training has become a standard approach for improving image and video generation, yet its extension to high-fidelity 3D generation remains challenging. Reward signals computed only from rendered images cannot fully characterize geometric quality, while direct supervision with explicit 3D metrics is ill-posed for single-image generation because unobserved regions have no uniquely determined geometry. We introduce Hybrid3D-RL, a hybrid-space post-training framework for VoxSet-based 3D generation. Its first component, the Implicit Latent Reward Model (I-LRM), reuses the geometric prior of a pre-trained generator to assess mesh quality directly from latent representations. We train this model without human annotation using preference pairs constructed through multi-granularity VAE re-encoding. Its second component, HRPO (Hybrid-space Ranking Preference Optimization), separates online rollout from reward evaluation in a modular training pipeline. HRPO combines an explicit appearance reward measured from a camera-corrected rendering with the implicit latent reward, and optimizes stochastic flow trajectories using relative advantages and reference-policy regularization. Experiments demonstrate reliable latent quality discrimination and consistent improvements in image-to-3D generation across quantitative and qualitative evaluations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.