acceptodds
Under review as a conference paper at ICLR 2027

Entity3R: An End-to-End Entity-Aware Geometric Foundation Model

Abstract

Class-agnostic 3D entity understanding aims to recover open-world entity identities with basic scene geometry from multi-view images or videos. Recent state-of-the-art methods built on instance-grounded geometric foundation models rely on hyperparameter-sensitive clustering to decode objects, which makes entity mask inference unreliable and costly. They also inherit the low native resolution of geometry transformers, limiting fine-grained entity perception. We present Entity3R, an end-to-end, entity-aware 3D/4D geometric foundation model. Entity3R extracts high-resolution geometry-instance representations with a dual-branch backbone, and decodes 3D/4D-consistent entity masks through learnable global object queries that interact with frame-level instance features. This design enables accurate and efficient 3D entity segmentation without post-hoc clustering. We further introduce a training-free, chunk-wise decoding strategy that allows Entity3R to decode more entities than its query budget and to process longer input sequences. To train Entity3R, we build EntityWorld, a large-scale corpus for joint geometry and instance learning, containing 360K sequences and 84.8M frames drawn from existing geometric datasets and newly collected photorealistic game data. Entity masklets are generated at scale by an automatic annotation pipeline. Trained on EntityWorld, Entity3R achieves state-of-the-art performance on both entity segmentation and geometric reconstruction benchmarks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.