acceptodds
Under review as a conference paper at ICLR 2027

GOLP: Global Object-Centric Learning via Spatial and Semantic Priors

Abstract

Grouping objects within an image does not automatically yield object knowledge that is reusable across scenes. This paper introduces GOLP, which uses spatial and category-related visual priors to learn global object representations for cross-scene recognition and compositional scene synthesis. A frozen SAM3 generates candidate masks with fixed dataset-level prompts, organizing object observations through their spatial extent and geometry. A frozen DINOv2 provides regional visual descriptors whose category-related structure guides shared prototype matching. The model retrieves content from a global object bank and separately encodes object states, background, and scene attributes, learning these factors through compositional feature reconstruction. In the second learning stage, the representation module and pretrained RAE are frozen, and only a Transformer adapter is trained for image decoding. Reconstruction, editing, and generation share the same latent variables and decoding pathway, without additional raw-image detail at decoding time. We evaluate cross-scene recognition, object discovery, and identity–state interventions, and provide additional analyses of mask supervision and prompt selection. The results examine how shared object content can be reused across scene conditions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.