Structured Voxel Diffusion for 3D Generation
Abstract
Recent 3D generative models often rely on multiview intermediates or learned latent representations, separating the generated variables from the final geometry. We introduce GRAVER, an image-conditioned framework that combines compact latent modeling of coarse structure with direct generation of fine geometry in an explicit sparse representation. Direct high-resolution geometric generation is challenging because volumetric cost grows cubically while the surface signal is highly sparse. GRAVER addresses this challenge with a staged block-UDF formulation that predicts (i) active blocks on a global grid, (ii) local surface-support masks within those blocks, and (iii) fine unsigned distance values only near the predicted support. The first two stages organize sparse structure, whereas the final stage predicts explicit UDF values without a geometry autoencoder. This decomposition concentrates computation near surfaces, accommodates open and non-watertight geometry, and reconstructs meshes through sparse marching cubes. The same spatially indexed representation also enables native local editing by preserving selected support and UDF values, and supports a preliminary adaptation to blockwise PBR prediction. Experiments show competitive image-shape alignment, validate the support-conditioned fine-UDF design, and demonstrate strong preservation of fixed regions during editing.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.