OpenU3R: Towards Zero-Shot Generalizable Feed-Forward Open-Vocabulary Understanding and 3D Reconstruction
Abstract
Jointly performing 3D Gaussian reconstruction and scene understanding in real-time is essential for embodied agents operating in unseen environments, yet existing feed-forward methods are restricted to closed vocabularies or struggle to generalize across datasets. Thus, we present OpenU3R, an open-vocabulary feed-forward framework that leverages complementary multi-view cues to support high-quality 3D Gaussian reconstruction together with semantic and instance-level scene understanding. Instead of using a fixed closed-set classifier, OpenU3R projects object queries into a shared vision-language embedding space and matches them against text embeddings, thereby endowing our method with open-vocabulary semantic capability. We further propose confidence-aware multi-view CLIP distillation, which pools frozen CLIP features within object masks across views and injects them into the object queries, improving cross-dataset transfer. On ScanNet, OpenU3R outperforms the best-performing baseline by 7.86 percentage points in mIoU, 5.43 percentage points in PQ, and 3.36 percentage points in mAP while maintaining competitive novel-view synthesis quality. Under zero-shot evaluation on RealEstate10K, OpenU3R improves over the same baseline by 29.00 percentage points in mIoU, 32.93 percentage points in PQ, and 11.95 percentage points in mAP, demonstrating strong cross-dataset generalization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.