CVE3D: Consistent Multi-View Expansion and Adaptive Fusion for Coherent and Reliable 3D Generation
Abstract
Insufficient observational constraints remain a major challenge for image-conditioned 3D generation, particularly in the single-image setting. A high-quality 3D generator can preserve visible evidence but must hallucinate the occluded side of an object. Novel-view diffusion provides complementary evidence, but independently generated views often drift in appearance, geometry, or pose. These cross-view inconsistencies impede the integration of view generation and multi-view-conditioned 3D generation into a unified pipeline. We present CVE3D, a consistent view-expansion framework that bridges this gap. A provisional 3D representation generated from the source image is rendered at each target camera pose to provide view-aligned RGB, depth, and surface-normal cues. Conditioned on these cues and the detached correction from the preceding view, our View-by-View Consistency Guidance Adapter corrects a selectively fine-tuned Zero-1-to-3 backbone to sequentially synthesize consistent target views, thereby reducing cross-view inconsistencies before 3D generation. We generate and rank complete trajectories based on source-referenced semantic consistency and adjacent-view continuity, rather than selecting each view independently. The selected views are subsequently integrated into TRELLIS.2 using our Consistency-Guided Adaptive Flow-Prediction Fusion, which guides cross-attention entropy fusion with view-consistency scores to limit the influence of residual unreliable observations during 3D generation. Experiments demonstrate that CVE3D achieves excellent overall generation quality, with consistently strong performance in semantic alignment and rendered-view fidelity. Qualitative comparisons further show detailed geometry and coherent appearance across viewpoints, confirming the effectiveness of our consistency-guided view expansion and fusion design. These results establish CVE3D as a practical and effective framework for integrating novel-view synthesis with multi-view-conditioned 3D generation. Code is available at https://anonymous.4open.science/r/CVE3D-AD3B.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.