SEG: Leveraging Self-Expanding Geometric State for Successive View Synthesis
Abstract
Recent advances in both reconstruction and multi-view diffusion have greatly improved the quality and geometric consistency of novel view synthesis (NVS). However, most existing methods either adopt joint multi-view synthesis for predetermined target views, or rely on video generation models for continuous synthesis, which limits their ability in successive view synthesis where incremental and unordered view queries are required. To address this limitation, we present SEG, a successive view synthesis framework built around a Self-Expanding Geometric state. The key insight of SEG is to turn context-dependent geometric latent into self-expanding states with consistency during expansion and causality under successive queries. Therefore, we construct our SEG state from early geometric encoder level based on consistency analysis, and design causal-prefix ground-truth state for supervision. We pair these latents with KV reference memory to form the final SEG state for efficient successive queries. Based on SEG state, we design our pipeline with a causal geometric diffusion model for state synthesis and expanding, and introduce a geometric adapter to map generated latents to multi-level features for RGB readout. We show that SEG state provides a well-suited solution for successive NVS, including extensive experiments on RealEstate10K and DL3DV with strong view synthesis quality, viewpoint accuracy, and computational efficiency.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.