A Contrastive Target Is a Demand, Not a Label
Abstract
An InfoNCE target is usually read as a label: it says which samples belong together. We show that it acts as a demand, and that what an encoder learns depends on how the demand can be met, not only on what it says. We write the target of a batch as a matrix with two parameters, sharpness and specificity, and build targets that agree exactly on class correctness, sharpness and the class composition of every batch, with the temperature tuned for every arm. The cleanest test compares two targets that name a random class-mate as each row’s positive and carry identical information. When the class-mate is fixed, encoders memorise the pairing. When it is redrawn at every step, they cannot, and held-out classes become tighter by 30 to 50% in angular spread and more accurate by 1.5 to 4.6 points of withinview class retrieval. Against the standard target, which names the true partner, the redrawn class-mate gains 3.4 and 3.3 points of class structure on two image settings and gives up 44 and 82 points of instance retrieval, a trade that reappears at larger scale on CIFAR-100 with a ResNet-18. The second parameter, sharpness, behaves as a second temperature. Over a 7 × 8 grid with every optimum interior, the best temperature falls 40× as sharpness grows 73×, the best smoothed target is within 2 points of the best-tempered one-hot target in three of four settings, and a closed form of the loss predicts the direction of this ridge. With the temperature free, sharpness is harmless on a correct target, and most of the damage a sharp wrong target does in a fixed-temperature loss disappears. Specificity, the axis the temperature cannot reach, is where the target’s content lives.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.