acceptodds
Under review as a conference paper at ICLR 2027

Native 3D Human Generation with Geometry-Aligned Soft Rigging from a Single Image

Abstract

Generating high-fidelity, riggable 3D human avatars from a single image remains a fundamental challenge. Native 3D models suffer from limited data and poor generalization. Meanwhile, 2D-inspired methods rely on multi-view or video supervision, which introduces cross-view inconsistency and unstable convergence. To address these limitations, we present NVWA, a native 3D generation framework that reconstructs geometry-aligned, riggable avatars from one image. Our method makes two key innovations. First, rather than supervising with multi-view renderings or Internet videos, we synthesize large-scale explicit 3D human assets for direct native 3D supervision. This eliminates the multi-view inconsistency inherent in distillation-based methods and enables substantial gains in geometric fidelity. Second, we depart from predefined parametric topologies (e.g., SMPL). Such templates constrain outputs to naked-body shapes and cannot model loose clothing or accessories. Instead, we propose soft rigging. Our model jointly generates free-form volumetric geometry and a spatially aligned semantic correspondence field. This field establishes per-vertex correspondence to a standard rigging template without restricting the generated shape. Shape generation is thus decoupled from template constraints while full pose controllability is preserved. To support training, we introduce a scalable pipeline for synthesizing photorealistic explicit 3D humans. We also release NVWA-100K, a dataset of 100,000 high-fidelity textured mesh models. Experiments demonstrate that NVWA achieves state-of-the-art fidelity and rigging quality, significantly advancing single-image 3D human generation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.