OISE: Objective-Induced Splits in Encoders
Abstract
What a frozen vision encoder splits images on first depends on whether its training objective matches a whole-image target or predicts hidden content. We examine this relationship in 21 frozen vision encoders, comprising 15 core models spanning objectives and architectures and six language-supervised variants spanning scale, pretraining data, a convolutional backbone, and a VLM tower. A fixed clustering procedure induces a binary partition within each class without group labels. We compare assignments on the same images using the Adjusted Rand Index (ARI), whose chance expectation is zero. On UrbanCars, 19 of the 21 encoders have self-consistency of 0.56–0.93 across discovery draws. Eight self-supervised cross-view encoders produce similar partitions (ARI 0.56–0.81). Label- and language-supervised models occupy an intermediate position, agreeing with the cross-view models at 0.37 on average. MAE, data2vec, I-JEPA, and the hybrid DINOv2 have low agreement with whole-image models (0.07) and with one another (0.06). The exception is MAE–data2vec (0.32 on UrbanCars and at most 0.08 elsewhere). Within the original four-dataset analysis, the three latent-prediction models (data2vec, I-JEPA, and DINOv2) are jointly isolated only on UrbanCars, where partitions are most stable. Matched training reproduces the ResNet-18 ordering with the same images, epochs, and batch size. Controlled interventions link partition selection to cross-view augmentation and show that relative variance can shift the partition among stably expressed factors. Finally, label-free partition agreement predicts cross-encoder group-transfer gains, although the gains are limited to a few worst-group accuracy points.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.