Representation Risk in Pretrained Image Encoders
Abstract
Applied researchers increasingly convert images into features with pretrained encoders, then use those features in a downstream prediction model. The encoder is often treated as an implementation detail. We show that it can instead be a consequential source of model uncertainty. We call this uncertainty representation risk: plausible pretrained encoders map the same images into different feature spaces and can yield sharply different out-of-sample conclusions. We compare modern and legacy frozen encoders across house prices, racehorse workout performance, breast-cancer histology, chest radiographs, facial age, and rice disease. With the same images, split, and linear-head protocol, replacing ResNet50 with DINOv2 raises test from 0.395 to 0.637 for houses and from 0.043 to 0.137 for horses. No encoder is best in every task. Validation-trained stacking improves further when strong encoders make complementary errors: a DINOv2–SigLIP 2 stack reaches 0.666 for houses, and averaging improves BreaKHis accuracy from 0.911 to 0.926. Stacking adds little for horses and facial age, and rice is at its performance ceiling. The rankings persist with nonlinear neural-network heads, so the gaps are not artifacts of linear extraction. These results support a simple workflow: benchmark plausible representations, select on locked validation data, combine only when validation evidence justifies the additional cost, and report paired uncertainty. We implement this workflow in a reproducible software framework using the "LookAgain-ML" package.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.