acceptodds
Under review as a conference paper at ICLR 2027

Select What Resists Compression: Reconstruction Code Length for Label-Free Data Selection

Abstract

Selecting training data typically requires labels, captions, or an externally pretrained encoder to assess data quality. A pool of unlabeled images offers none of these, and what makes an image useful for self-supervised pretraining is itself unsettled. We show that the per-image loss of a corruption–reconstruction network trained on the pool is a conditional code length: the number of nats needed to encode the image given a masked and noised view of it. We rank the pool by this quantity, which we call the CORE score. In other words, we keep the images that a model of the pool compresses worst. On ImageNet-1k, MAE pretrained on the top-scored 50% at half the gradient steps reaches 62.1% under linear probing against 62.3% for the full set, whereas a random 50% reaches 56.7%; at equal steps the top-scored half exceeds the full set by 5.9 linear-probe points and 1.1 ADE20K mIoU, and a random half by 3.3 and 0.8. The ordering top >random > bottom holds under fine-tuning for five pixel-prediction backbones and for DINO, which never reconstructs pixels, and a scorer trained on half of the pool ranks the other half almost identically (ρs = 0.89). A 1.28M-image subset selected from ImageNet-21k exceeds the human-curated ImageNet-1k by 0.32 points on average. The score correlates with lossless compressed size (ρ = 0.61), yet the lowest-scored images include high-frequency repetitive textures: heavy masking discounts the bits that a generic compressor counts. We report the one-time cost of the scorer together with the number of pretraining runs after which it pays off. Code is provided in the supplementary material.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.