Towards the Scaling Law of Neural Model Probes
Abstract
Probing, the practice of training a lightweight classifier on frozen model representations and reading off its accuracy, has become a popular tool for studying what neural network models encode. Yet probes themselves have received little systematic study: design choices such as architecture, hidden width, and initialization are often made arbitrarily, and probing accuracy alone cannot separate information present in the representation from what the probe learns on its own. In this paper, we study the science of probes, and ask whether their behaviors follow scaling laws analogous to those of model pretraining. We use minimum description length (MDL; Voita & Titov, 2020) as our measure of probe faithfulness, and study its relationship with complexity of MLP probes (i.e., hidden width) and their training data size across data modalities, analyzed models, tasks, layers, and initialization schemes. We find that (1) faithfulness is non-monotone in probe width, exhibiting a characteristic U-shape with an optimum typically between widths of 64 and 256; (2) this optimum can be predicted from simple representation-level features, namely the layer depth and the accuracy of a linear probe; (3) dataset size affects faithfulness monotonically and exponentially; and (4) the scale of probe initialization shifts the optimum in a log-linear fashion, while maximal update parametrization flattens the faithfulness–complexity curve. Together, these findings set the first steps toward establishing a scaling law of probing and offer practical guidance for designing and configuring probes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.