acceptodds
Under review as a conference paper at ICLR 2027

DVSM: The Simplest Decoder-Only View Synthesis Model Performs Best

Abstract

Regression-based view-synthesis models have explored encoder-decoder, decoder-only, and foundation-model designs, but it remains unclear which composition is most effective. We revisit this design space through controlled experiments and find that the simplest choice works best: one decoder should both reconstruct a reusable scene state and render novel views. DVSM prefills a per-layer KV cache from the context views, then renders each target through unidirectional cross-attention using the same weights. Every tested form of weight decoupling reduces quality despite increasing the parameter count. Our feature-space analysis provides a consistent explanation: for the same viewpoints, complete weight sharing keeps reconstruction and camera-only rendering features aligned and enables the renderer to recover the corresponding oracle features. Foundation features can be injected during context prefill as an orthogonal enhancement. DVSM establishes state-of-the-art results on DL3DV and zero-shot out-of-distribution benchmarks, outperforming both geometry-free and geometry-based baselines without explicit 3D inductive biases. On most benchmarks, our best-performing variant is also the smallest model, showing that the gains come from our design rather than model scale.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.