acceptodds
Under review as a conference paper at ICLR 2027

Vid2Orbit: Adapting Video Diffusion as Multi-View Diffusion for Single-Image Scene Reconstruction

Abstract

Video diffusion models provide strong priors for single-image 3D reconstruction, but their temporal representation is designed for dense video rather than sparse reconstruction views. In the native video VAE, the reset frame follows a higher-fidelity single-frame path, whereas subsequent frames are temporally compressed; under viewpoint motion, these neighboring frames provide limited new angular coverage while carrying larger codec errors. We verify this effect without diffusion sampling: resetting the frozen VAE for every ground-truth view improves round-trip PSNR by 6.25 dB at view spacing. Motivated by this finding, we introduce Vid2Orbit, which repurposes a pretrained video diffusion model as a sparse multi-view generator. Each requested camera uses the VAE's native single-frame path, while the diffusion transformer jointly generates the complete view set. We further introduce camera-aligned feature exchange: intermediate multi-view features are aggregated in a shared 3D field and read back along camera rays. The resulting camera-indexed views provide a practical interface to COLMAP and Gaussian reconstruction.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.