acceptodds
Under review as a conference paper at ICLR 2027

DreamSurfels: Single-Image 3D Generation with Geometry-Conditioned Video Diffusion

Abstract

Generating a detailed 3D mesh from a single image requires both plausible completion of unseen regions and consistent 3D geometry across the generated views. Existing methods that reconstruct 3D from synthesized views can propagate inconsistencies across viewpoints into the final mesh. Video diffusion provides a powerful prior for coherent view sequences, yet visual coherence does not guarantee that the views share a consistent 3D surface. To this end, we introduce DreamSurfels, a geometry-conditioned video diffusion framework for single-image 3D mesh generation. To provide explicit geometric guidance to video diffusion, we lift the visible surface from the input view into Gaussian surfels and render the partial geometry at target viewpoints. Residual conditioning branches inject these renderings into the diffusion backbone, anchoring the generated views to the observed surface as the model completes unseen regions. Beyond consistent view synthesis, accurate mesh reconstruction requires surface orientation cues that RGB appearance alone cannot reliably provide. We therefore develop a joint RGB-normal video diffusion model that uses intermediate RGB features to guide surface normal predictions, aligning geometry with the generated views. With the resulting geometrically consistent RGB views and surface normals, our framework then reconstructs detailed textured meshes. Experiments show that DreamSurfels produces consistent multi-view geometry and high-quality mesh reconstructions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.