acceptodds
Under review as a conference paper at ICLR 2027

FlashLVSM: A Constant-Time View-Synthesis Transformer

Abstract

Latent scene representations have emerged as a promising alternative for novel view synthesis, encoding multi-view observations into learned features without relying on predefined geometric primitives or rendering operations. Despite their promise, realizing them with modern neural architectures, such as transformers, often incurs substantial rendering cost. As the number and resolution of source views increase, every target token repeatedly attends to a growing set of source tokens through a deep neural renderer, causing the rendering cost to grow rapidly. We present FlashLVSM a constant-time view-synthesis model designed for scalable, highly efficient rendering from latent scene representations, which is carefully co-designed around four complementary principles, each addressing a distinct source of rendering cost. A richer, more 3D-expressive latent representation enables us to use a lightweight decoder without sacrificing rendering quality. Target-aware view selection restricts computation to source views geometrically relevant to the target, while target-aware token pruning further retrieves only the informative tokens within them, making rendering cost independent of the total number of input views. Finally, mixed-precision inference accelerates the remaining decoder computation. Together, FlashLVSM renders at 142 FPS with 32 input views on a single consumer-grade GPU, and maintains roughly 140 FPS as the number of views increases to 256 while preserving rendering quality. This corresponds to speedups of nearly 12x and 95x, respectively, with 204x fewer FLOPs and 25x lower peak memory at 256 views. Despite this reduction in computation, FlashLVSM remains competitive in rendering quality, matching or exceeding the baseline from 64 input views onward. These results establish efficient latent scene representations as a compelling alternative to conventional explicit representations for efficient novel view synthesis.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.