Select What Resists Compression: Reconstruction Code Length for Label-Free Data Selection
Abstract
Selecting training data typically requires labels, captions, or an externally pretrained encoder to assess data quality. A pool of unlabeled images offers none of these, and what makes an image useful for self-supervised pretraining is itself unsettled. We show that the per-image loss of a corruption–reconstruction network trained on the pool is a conditional code length: the number of nats needed to encode the image given a masked and noised view of it. We rank the pool by this quantity, which we call the CORE score. In other words, we keep the images that a model of the pool compresses worst. On ImageNet-1k, MAE pretrained on the top-scored 50% at half the gradient steps reaches 62.1% under linear probing against 62.3% for the full set, whereas a random 50% reaches 56.7%; at equal steps the top-scored half exceeds the full set by 5.9 linear-probe points and 1.1 ADE20K mIoU, and a random half by 3.3 and 0.8. The ordering top >random > bottom holds under fine-tuning for five pixel-prediction backbones and for DINO, which never reconstructs pixels, and a scorer trained on half of the pool ranks the other half almost identically (ρs = 0.89). A 1.28M-image subset selected from ImageNet-21k exceeds the human-curated ImageNet-1k by 0.32 points on average. The score correlates with lossless compressed size (ρ = 0.61), yet the lowest-scored images include high-frequency repetitive textures: heavy masking discounts the bits that a generic compressor counts. We report the one-time cost of the scorer together with the number of pretraining runs after which it pays off. Code is provided in the supplementary material.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.