Endo4DLR: Online Language-Aligned 4D Reconstruction with Persistent Gaussian Anchors
Abstract
Dynamic intraoperative scene understanding requires jointly recovering evolving 3D geometry and grounding language semantics in reconstructed scenes. Yet performing both tasks online from unposed endoscopic streams remains challenging due to camera motion, tissue deformation, and changing visibility. Existing methods either rely on per-scene optimization or focus primarily on geometric reconstruction, leaving this joint problem largely unexplored. In this work, we introduce , a fully feed-forward framework for online language-aligned 4D reconstruction that jointly predicts camera parameters and language-embedded Gaussians without per-scene optimization. Our key idea is to maintain persistent Gaussian anchors that connect current local observations with reusable historical scene information. In particular, we introduce adaptive observation queries that capture fine-grained details of the current frame and are organized into local 3D Gaussian groups. In addition, we introduce spatial anchors that retrieve and aggregate historical features to predict Gaussian attributes and language embeddings. By combining current geometric estimates with retrieved historical context, Endo4DLR adapts to scene evolution while maintaining temporal consistency in reconstruction and language alignment. Extensive experiments on multiple endoscopic datasets demonstrate that Endo4DLR outperforms state-of-the-art online methods in reconstruction, camera estimation, and language grounding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.