Transformers Align in Subspaces, Not Units: Rotation-Aware Model Merging
Abstract
Model merging combines trained networks in weight space, and every alignment method implicitly picks a symmetry group for comparing them. We show that group is mis-specified for Transformers. Transport-based alignment searches the Birkhoff polytope ; self-attention is also invariant under continuous rotations of its query-key and value-output subspaces, and is exactly the permutations, so no transport plan - however soft - expresses a non-trivial rotation. The LayerNorm residual stream, in turn, admits among isometries exactly the rotations fixing the mean axis, where a soft plan is inexact, not merely suboptimal. Rotation-aware Optimal Transport (RaOT) aligns in that quotient. Building on interface-wise rotation fusion (Zhang et al., 2025), it differs in four ways: activation-estimated alignments from one forward pass; head correspondence and within-head rotation solved jointly as one exact assignment; a mean-preserving Procrustes map on the residual stream; and identity competing with every estimate, which supplies a provenance switch and a test of whether alignment should act at all. This is a mechanism paper, not a leaderboard entry. No method - ours included - merges independently trained ImageNet ViTs to usable accuracy, so we read the barrier through matched contrasts: across nine public pairs and a four-width series inside one model family, what the isometries' estimator adds beyond the permutation's is larger at larger width. The polytope's own share shrinks only on the original three-pair ladder - across pairs it straddles zero, and a pre-registered same-recipe study reverses it - so we claim the isometries half alone, on recipe- or corpus-different parents, and not a width law. Alignment buys accuracy in frozen-feature transfer and, for parents differing in task and initialization, in trunk-head compatibility rather than in the backbone representation. On fine-tuned siblings of one initialization the gate rejects every transform: the role there is detection, not repair.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.