acceptodds
Under review as a conference paper at ICLR 2027

DeOR: Object-centric Representations from and for Multimodal LLMs

Abstract

Representing objects individually supports scene understanding and object-level manipulation. Slot-based methods decompose scenes into object representations, but many rely on a fixed number of slots despite varying object counts across images. Moreover, these representations often lack explicit object descriptions, making it difficult to identify and select individual objects for manipulation. To address these limitations, we propose DeOR (Description-guided Object Representations), a framework that uses a multimodal large language model (MLLM) to construct a variable number of object representations that can be identified through language. DeOR sequentially generates an object description and a corresponding object token for each object. Each description specifies which object its paired token should represent. Each representation combines the token’s semantic information with visual features from the image. This structure allows objects to be identified through their descriptions and manipulated individually through their representations. DeOR supports compositional generation by recombining object representations across images. Its generated descriptions also allow an external language model to select objects based on user instructions. DeOR then edits the selected representations to remove objects or transfer appearance from another image. On COCO scenes with 4-6 objects, DeOR-L achieves a 26.6% relative improvement in FG-ARI over the strongest evaluated baseline. In zero-shot transfer from COCO to PASCAL VOC, it achieves a 44.2% relative improvement in instance-level mBO over the strongest VOC-trained baseline.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.