acceptodds
Under review as a conference paper at ICLR 2027

Convergent Evolution of Neural Networks: Zippering and Weak-Strong Equivalence

Abstract

Why and when should independently trained neural networks develop the same internal mechanisms? We develop a theory of this convergent evolution through *identifiability*: what a network's end-to-end function determines about its internal computation. Our *zippering* theorems establish when agreement at the end forces corresponding representations throughout a nonlinear network, up to linear transformations. Our *weak-strong equivalence* theorems then show when this correspondence forces individual channels and gates to recur. Tasks that require a greater fraction of these components force stronger convergence. We prove exact and approximate versions of both results. Most strikingly, when our theory is applied to modern transformer architectures, it predicts *privileged attention heads*: individual head computations that recur across models. We find privileged heads and gates in both vision and language, ordered alignment across unequal depths, and stronger recurrence of more-used components. Task constraints can therefore force a shared computational hierarchy and reproducible mechanisms across models trained on the same tasks. These results thus give mechanistic interpretability stable structures to explain and AI safety research recurring targets for intervention in models spanning different tasks, modalities, and architectures.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.