acceptodds
Under review as a conference paper at ICLR 2027

From 3D Representations to Object Tokens in Object-Centric MLLMs

Abstract

Multimodal large language models (MLLMs) provide a general interface for perception and language-based reasoning, yet it remains unclear how the underlying 3D representation affects the performance. Existing systems often couple object discovery, representation source, and subsequent object-token construction, making their individual contributions difficult to isolate. In this paper, we study these factors under a controlled object-centric formulation that separates object proposals from the representation used for downstream reasoning, and compare multiple representation sources and object-token construction mechanisms. Under our strongest configuration, the resulting model surpasses prior point-cloud-only methods on most comparable metrics across five ScanNet benchmarks. Our experiments further show that object discovery and object representation can be treated as distinct design choices, and that the choice of frozen 3D representation has a substantial effect on downstream performance, whereas increasing the complexity of the object-token construction mechanism provides only limited and task-dependent gains. These results highlight representation source as a central design factor in object-centric 3D MLLMs and show that effective object representations can be constructed with comparatively lightweight aggregation mechanisms.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.