T-LRM: Temporal-aware LRM for multi-view dynamic scene reconstruction
Abstract
Large Reconstruction Models (LRMs) have recently emerged as the dominant paradigm for feed-forward 3D reconstruction. However, existing 3D foundation models are primarily designed for static scenes or monocular dynamic inputs, making them difficult to extend to high-quality multi-view dynamic scene reconstruction. A straightforward frame-by-frame inference strategy is computationally expensive and fails to exploit the strong temporal redundancy inherent in multi-view videos. To address these limitations, we propose T-LRM, a Temporal-aware LRM framework that extends existing 3D foundation models to temporally aware multi-view dynamic reconstruction in a training-free manner. Our key insight is that multi-view dynamic videos contain substantial spatial redundancy across time. By identifying dynamic regions while sharing temporally consistent static regions, redundant computation can be largely eliminated and cross-temporal interference can be effectively suppressed. Based on this observation, we introduce a Temporal-aware Token Merging strategy that anchors the scene geometry using shared static tokens while reconstructing dynamic objects with timestamp-specific dynamic tokens. The proposed framework is lightweight, plug-and-play, and can be seamlessly integrated into different feed-forward reconstruction architectures without retraining. Extensive experiments on multiple benchmark datasets demonstrate that T-LRM achieves multi-fold inference acceleration while maintaining comparable reconstruction quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.