Iterate, Don't Stack: Grounded Recursive Transformers for Generalizable Novel View Synthesis
Abstract
Existing feed-forward novel view synthesis transformers apply a fixed amount of computation to render each target view per scene, although the complexity within a single scene is not distributed evenly. By fixing the number of transformer layers, they freeze the trade-off between quality and computation. We propose Universal View Synthesis Model, UVSM, capable to adaptively allocate computation within the target image. UVSM follows a unique principle: ground the scene once and recursively refine each token corresponding to image patches in parallel until convergence. A halting probability gate acting on refined tokens determines whether they need further refinement step. The halting gate acts as a dial at inference between a fast preview and a converged render. One trained model therefore replaces a family of fixed-depth models that a quality-compute curve would otherwise require. We further show that increasing the number of refinement steps in our network generalizes better than expanding the depth of a transformer network. Under the same training protocol, UVSM beats every baseline on all seven benchmarks with the highest margins on the out-of-distribution evaluations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.