AttenDiffCrafter: Attention-Crafted Consistent Novel Views from a Single Image
Abstract
Generating consistent novel views from a single image is challenging under large camera motions, where unseen regions must be synthesized while preserving scene structure across frames. Existing approaches face different limitations: autoregressive Transformers tend to accumulate errors over long sequences, whereas diffusion models built on U-Net backbones have limited global interaction and may produce color shifts or implausible content in regions outside the input view. We propose AttenDiffCrafter, a diffusion framework for novel-view synthesis built on attentioncentric Transformer architecture. Instead of relying on convolution-dominated feature extraction, the model uses self-attention throughout as the primary mechanisms for spatial and cross-view feature interaction. To better relate different viewpoints, we introduce a learnable query-key cross-attention module that learns view-dependent correspondences between source and target features. Camera motion is represented by relative poses and Plücker coordinates, and temporal attention is used to model the relations along the camera trajectory. Experiments on Matterport3D and RealEstate10K show that AttenDiffCrafter achieved a good overall performance in terms of visual quality, geometric consistency, and temporal coherence.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.