acceptodds
Under review as a conference paper at ICLR 2027

SlotMask: Masked Feature Reconstruction for Object-Centric Representation Learning

Abstract

Humans perceive the world as a composition of distinct objects. Inspired by this, object-centric learning represents scenes as sets of object-level vectors called slots. Recent approaches learn slots by reconstructing pretrained visual features with slot-conditioned autoregressive transformers. We introduce SlotMask, which replaces causal reconstruction with a bidirectional transformer trained to predict randomly masked features. Central to our approach is tying masked reconstruction to the object structure discovered by Slot Attention: each masked token is initialized with a mixture of slots weighted by its Slot Attention assignments, and decoder cross-attention is trained to align with these assignments. Together with a randomly sampled masking ratio, these mechanisms make masked reconstruction an effective objective for learning object-centric representations. SlotMask achieves state-of-the-art unsupervised object discovery on PASCAL VOC, MS COCO, and MOVi-C/E, outperforming autoregressive baselines under matched backbones, same training budgets, and with cheaper decoding.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.