acceptodds
Under review as a conference paper at ICLR 2027

Permutation Alignment and Depth in Weight-Tied Language Models

Abstract

Averaging the weights of two independently trained language models can yield a far worse model, even after their hidden units are permuted into correspondence. We ask how much of this damage permutation alignment removes, and how to interpret the remaining damage in weight-tied (looped) models, whose depth can vary at runtime. We measure fifteen pairs of 35M-parameter looped models at 16 and 192 loops, before and after a permutation alignment that respects the shared weights, recording the quality of the averaged (midpoint) model and both endpoints. Before alignment, the median barrier (the ratio of midpoint to better-endpoint perplexity) is 1.8 dex (base-10 log units) higher for eight pairs trained mostly with random loop counts than for six trained at a fixed 16 loops; after alignment the gap is under 0.05 dex. A shared residual remains: with the best of four alignment methods, midpoint perplexity is at least 12 times the better endpoint's on every pair. Equal-norm random perturbations reproduce its typical size; perturbations along the aligned weight difference's leading directions cause more damage than orthogonal ones on every pair tested. At 192 loops, endpoint quality sets the barrier: it falls where endpoints become unusable and persists between depth-stable ones. Continuing training from the aligned midpoint for 30% of the original budget removes over 90% of its residual on two depth-stable pairs, and distilling a probability-mixture teacher keeps deep-runtime degradation below 0.05 dex, versus 0.40 dex for plain continuation. Relative barriers should thus be reported with absolute endpoint quality and runtime depth.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.