How Representations Shape Label Requirements
Abstract
Pretraining can reduce the need for labels, but what makes a representation label-efficient? Labels reveal both what to predict and which inputs should share a prediction. We isolate these roles in a regression family whose candidate groupings have equal response complexity. Choosing a grouping before labels arrive can make a target risk unattainable, regardless of label count. A frozen representation retaining both candidates allows a supervised readout to overcome this structural barrier, at a higher estimation cost. We give upper and lower bounds on optimal risk within the committed class that match within a factor of three, and computable label budgets for known Gaussian tasks. Experiments on electrocardiograms, ImageNet, and robot actions vary input coverage and encoder update permissions. Under fixed recipes and on the measured grids, adaptation reaches specified mean-risk targets with one-tenth as many labels on ImageNet and robot actions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.