acceptodds
Under review as a conference paper at ICLR 2027

DOES IMAGE QUALITY AWARENESS EMERGE FOR FREE FROM SELF-SUPERVISED PRETRAINING?

Abstract

Vision foundation models are trained to recognize and describe image content, never to judge image quality. We ask whether they nevertheless encode humanperceived quality, and where in the network. We freeze four base-size Vision Transformers (ViT-B) trained with different objectives — masked-image reconstruction (masked autoencoder, MAE), self-distillation (DINOv2), image–text contrastive learning (Contrastive Language–Image Pre-training, CLIP), and supervised classification — and fit only linear probes to human quality ratings on two image quality assessment (IQA) benchmarks, one with authentic (KonIQ-10k) and one with synthetic (KADID-10k) distortions. All four encode substantial quality information. On authentic distortions their output features outperform a probe on hand-crafted low-level statistics, reaching a Spearman rank correlation (SROCC) of up to 0.87; within a single scene they rank synthetic distortions largely in line with human ratings; and they transfer between the two kinds of distortion, where the hand-crafted baseline barely does. The information is richest in intermediate blocks, where probes reach 0.92, and how much of it survives to the output depends on the objective: pixel reconstruction preserves nearly all of it, while discriminative objectives discard part of it. Hand-crafted statistics still identify the distortion type better, and we find no reliable link between quality encoding and robustness to corruptions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.