acceptodds
Under review as a conference paper at ICLR 2027

Revisiting Jigsaw Objectives for Fine-Grained Visual Representations

Abstract

Current self-supervised models focus on semantic invariance, capturing similarity of content while ignoring fine-grained variations in the input. This is useful for tasks such as semantic segmentation and retrieval, but physical scene understanding, anomaly detection, and the guidance and evaluation of generative models require representations that capture fine details and spatial structure of the input. This motivates unsupervised objectives that teach models more about the structure of what they encode. Jigsaw-style objectives are a natural candidate. Decomposing an image and asking a model to reassemble it requires attending to boundaries and local structure, and the supervision comes entirely from the decomposition. Such objectives were explored in early self-supervised learning but largely abandoned after contrastive methods proved stronger on classification, and they have not been revisited at the scale or in the architectures now standard. In this paper, we revisit this family in a modern self-distillation framework. We study a range of formulations of the Jigsaw puzzle objective, several of which are new, and evaluate them on several benchmarks. Across tasks requiring fine detail, including object discovery, tracking, and anomaly detection, these objectives outperform prior self-supervised models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.