GOLP: Global Object-Centric Learning via Spatial and Semantic Priors
Abstract
Grouping objects within an image does not automatically yield object knowledge that is reusable across scenes. This paper introduces GOLP, which uses spatial and category-related visual priors to learn global object representations for cross-scene recognition and compositional scene synthesis. A frozen SAM3 generates candidate masks with fixed dataset-level prompts, organizing object observations through their spatial extent and geometry. A frozen DINOv2 provides regional visual descriptors whose category-related structure guides shared prototype matching. The model retrieves content from a global object bank and separately encodes object states, background, and scene attributes, learning these factors through compositional feature reconstruction. In the second learning stage, the representation module and pretrained RAE are frozen, and only a Transformer adapter is trained for image decoding. Reconstruction, editing, and generation share the same latent variables and decoding pathway, without additional raw-image detail at decoding time. We evaluate cross-scene recognition, object discovery, and identity–state interventions, and provide additional analyses of mask supervision and prompt selection. The results examine how shared object content can be reused across scene conditions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.