SUPRA-Vox: Native Voxel Prompts for Super-Aligned 3D Asset Generation
Abstract
Image-to-3D models can generate visually compelling assets, but often struggle to faithfully reproduce the geometry visible in the input view. In multi-stage 3D generation, missing parts or misplaced surfaces in the initial structure can persist in the final shape, even after later stages add fine details. We propose SUPRA-Vox, a surface-guided framework for sparse voxel generation toward what you see is what you get geometry. SUPRA-Vox encodes the visible surface reconstructed by a geometry foundation model intoNative Voxel Prompts using the generator's pretrained variational autoencoder (VAE). These prompts remain available as separate conditions throughout denoising. Together with a visual hull mask, they guide the inpainting transformer as it jointly updates the complete voxel structure. The resulting structure is passed to a frozen downstream generator for detailed shape synthesis. We train SUPRA-Vox on a curated collection of high-quality 3D assets and evaluate it on two complex-geometry benchmarks we construct, Complex-Geo and DTC-Geo. Experiments on these benchmarks demonstrate state-of-the-art visible-surface reconstruction performance among the evaluated methods and improved overall geometric quality over native Pixal3D, demonstrating that better early structure benefits final geometry.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.