Rethinking Large View Synthesis Models: Capacity and Scene Memory
Abstract
Large view synthesis models (LVSMs) employ diverse architectures and scene representations, but how to allocate capacity and computation for effective scaling remains unclear. We organize these systems through scene memory constructed from context views and queried for novel view synthesis. We systematically study network organization, spatial granularity, model capacity, and the allocation of re- sources between scene memory and decoding. Within the evaluated full-attention configurations, decoder-only models favor reconstruction quality, whereas sepa- rate decoders favor rendering efficiency. Spatial granularity is a key scaling di- mension alongside model capacity, and the benefits of larger scene memory de- pend on its decoder. These findings lead us to select and train a ViT-L/16 back- bone, yielding highly competitive reconstruction quality with shorter reported re- construction times than most baselines on DL3DV. Using this backbone, we ana- lyze sparse attention for novel view synthesis. Training-free block-sparse attention preserves spatial tokens and scene-memory capacity, accelerating scene construc- tion by 1.51–2.22× over dense inference at fixed inputs with only a small quality loss. Further analysis shows how layer sensitivity and routing overhead govern the quality and speed gains. We hope our study helps the community scale up LVSMs more effectively and explore better model designs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.