GRPQ: Structure-Preserving Alignment of 3D Diffusion Models with RGB-Only Rewards
Abstract
Aligning 3D diffusion models with specialized styles and human preferences is essential for producing usable assets, yet directly specifying such objectives in 3D remains difficult and costly. Generic 2D reward models provide abundant supervision, but optimizing their scores over rendered views can rapidly erode the structural prior of a pretrained 3D model. Existing safeguards often rely on structure-aware multi-view models or auxiliary depth and normal estimators, thereby limiting the range of applicable 2D rewards and potentially introducing estimator-specific biases. Our analysis reveals that pronounced disagreement between the cross-view rankings induced by 2D rewards and structural evidence often results in unreliable supervision and structural degradation; we refer to this phenomenon as reward–structural-evidence mismatch. Based on this finding, we introduce GRPQ, a prior-preserving alignment framework that uses the degree of mismatch as evidence for estimating the reliability of each view-level reward and assigns optimization weights accordingly. By suppressing harmful signals rather than introducing additional structural supervision, GRPQ enables generic RGB-only rewards to improve preference alignment while preserving the pretrained 3D structural prior. Extensive experiments across three alignment methods, multiple 3D diffusion models, and diverse reward models show that GRPQ improves structural integrity and prior retention while largely preserving reward gains and training efficiency. Process-level visualizations further show that GRPQ prevents optimization from favoring views that conceal geometric defects over those that reveal them, providing interpretable evidence for the proposed mechanism.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.