acceptodds
Under review as a conference paper at ICLR 2027

Rank Is Not Quality: Label-Free Model Selection Inverts under Instance-Private Nuisance

Abstract

Self-supervised encoders are increasingly selected without labels, most often by the effective rank of their embeddings (RankMe), on the premise that a more spread-out embedding is a more useful one; anti-collapse objectives such as VICReg and LeJEPA push the same quantity up during training. Because such criteria are used where labels are absent, a failure goes unnoticed unless labels are brought back. We show that effective rank can rank models backwards. When a label-irrelevant, instance-specific panel is appended beside each CIFAR-10 image, effective rank rises (9/9 paired runs) while linear-probe accuracy falls from to . Across that population RankMe's correlation with accuracy is (), it inverts within each of SimCLR, VICReg and BYOL, and a preregistered 30-run replication reproduces the inversion before convergence ( at epoch 5). A first proposition shows why the objectives cannot rule this out: an encoder that whitens (and, where required, Gaussianises) instance-private content attains the alignment optimum and the optimum of the marginal health terms in use, so no such objective can prefer semantic content. Nor is the failure specific to rank. On a preregistered dial from shared to instance-private nuisance, every criterion we measured prefers a nuisance-bearing model to the clean one somewhere: spread criteria from mid entropy upward, cluster criteria such as CLID's cluster learnability wherever the nuisance is shared; a second proposition shows that no criterion reading only the embedding distribution can be guaranteed safe. Three of the dial's five preregistered tests failed, as did an anti-collapse term designed to fix the problem; we report them all, claim no graded inversion, and restricted our claim to instance-private nuisance only after seeing them. The accuracy collapse itself is Chen et al.'s RandBit result; its consequence for label-free selection is, to our knowledge, new. Where candidate models may differ in what they capture, rank should not select among them, and no distribution-only criterion replaces a labelled check.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.