Convergent Evolution of Neural Networks: Zippering and Weak-Strong Equivalence
Abstract
Why and when should independently trained neural networks develop the same internal mechanisms? We develop a theory of this convergent evolution through *identifiability*: what a network's end-to-end function determines about its internal computation. Our *zippering* theorems establish when agreement at the end forces corresponding representations throughout a nonlinear network, up to linear transformations. Our *weak-strong equivalence* theorems then show when this correspondence forces individual channels and gates to recur. Tasks that require a greater fraction of these components force stronger convergence. We prove exact and approximate versions of both results. Most strikingly, when our theory is applied to modern transformer architectures, it predicts *privileged attention heads*: individual head computations that recur across models. We find privileged heads and gates in both vision and language, ordered alignment across unequal depths, and stronger recurrence of more-used components. Task constraints can therefore force a shared computational hierarchy and reproducible mechanisms across models trained on the same tasks. These results thus give mechanistic interpretability stable structures to explain and AI safety research recurring targets for intervention in models spanning different tasks, modalities, and architectures.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.