V23D: LEVERAGING VIDEO DIFFUSION PRIORS FOR HIGH-FIDELITY MULTI-VIEW 3D GENERATION
Abstract
Single-image 3D generation must infer self-occluded regions from incomplete evidence. Orbit videos offer complementary views, but a shared keyframe subset can omit cues essential to individual regions. We present V23D, a two-phase videoprior-to-3D framework that combines geometry-aware orbit generation with regionspecific conditioning. In Phase I, we adapt Wan2.2 with noise-decoupled LoRA and a Geometry-aware Orbit Adapter (GOA) to generate 49-frame turntable priors, reducing FVD from 205.8 to 176.8. In Phase II, we introduce a Regional Source Adapter (RSA) that connects the complete orbit to coarse-to-fine 3D generation. The sparse-structure stage records regional source-frame indices and their original attention masses; the shape stage uses independent current queries to re-read the referenced patches, while receiving sparse support separately. This design preserves source access across generation stages rather than passing only fixed fused features. On HY3D-Bench-test, V23D achieves 0.2418 ULIP-I and 0.3827 Uni3D-I, outperforming the evaluated baselines. Compared with the fourview OFA configuration, RSA improves F-score@1% from 57.4% to 61.1%, and achieves 62.7% on Toys4K in cross-dataset evaluation. Matched full-orbit comparisons further favor RSA over all-frame aggregation, instantaneous regional routing, and fixed-feature transfer, supporting the value of cross-stage source preservation and state-dependent patch re-reading.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.