DVSM: The Simplest Decoder-Only View Synthesis Model Performs Best
Abstract
Regression-based view-synthesis models have explored encoder-decoder, decoder-only, and foundation-model designs, but it remains unclear which composition is most effective. We revisit this design space through controlled experiments and find that the simplest choice works best: one decoder should both reconstruct a reusable scene state and render novel views. DVSM prefills a per-layer KV cache from the context views, then renders each target through unidirectional cross-attention using the same weights. Every tested form of weight decoupling reduces quality despite increasing the parameter count. Our feature-space analysis provides a consistent explanation: for the same viewpoints, complete weight sharing keeps reconstruction and camera-only rendering features aligned and enables the renderer to recover the corresponding oracle features. Foundation features can be injected during context prefill as an orthogonal enhancement. DVSM establishes state-of-the-art results on DL3DV and zero-shot out-of-distribution benchmarks, outperforming both geometry-free and geometry-based baselines without explicit 3D inductive biases. On most benchmarks, our best-performing variant is also the smallest model, showing that the gains come from our design rather than model scale.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.