Correspondence over Content: What Makes Task Information Readable in Neural Network Weights
Abstract
Weight space methods read trained networks to retrieve, compare and edit them. We ask what makes the information in a network's weights readable, and what it costs to read. Studying collections of networks that differ only in how a fixed set of images is labelled, so that the task is the only variable, we find that correspondence between models governs not what a reader can recover but what recovering it costs. When fixed sparse wiring pins each unit to the same inputs in every network, a linear probe identifies the task from a few hundred first layer weights or fewer, with no net growth over a sixteenfold range of width. Dense networks hold the same information in linearly readable form, but only in aggregates over their hidden units: the cheapest reader we found must read every unit, and its cost grows with the layer, with an exponent between and . Matching units to a template, on a few unlabelled images or on the weights alone, brings the dense budget back to a few hundred coordinates, so the premium for missing correspondence itself grows with the network, from threefold to seventeenfold across our widths. In the hidden to hidden layer, removing the permutation symmetry with fixed wiring is not enough: its coordinates stay unreadable, so removing a symmetry is not the same as supplying correspondence. Along the way we document how standard readings mislead: per model scale alone identifies the training dataset on a public model zoo, and an alignment step can itself carry the task.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.