acceptodds
Under review as a conference paper at ICLR 2027

Pairwise Anti-Collapse Regularization Restores Representational Diversity in Pure-Attention Transformers

Abstract

Pure self-attention without skip connections converges to a rank-one subspace at a doubly exponential rate in depth—a pathology known as representational collapse. For a decade the remedy has been architectural: route around attention, treating collapse as an inevitability to be engineered against. We show it is instead learned: the representation begins near full rank and contracts as training proceeds, because cross-entropy finds that aligning tokens cheaply reduces loss. From one initialisation, two gradients therefore steer the same network to opposite fixed points—the rank-one corner, or an orthonormal frame. Reaching the frame means penalising the right object: not the channel covariance of feature-decorrelation methods, since decorrelated channels do not stop two tokens from sharing a direction, but the token cosine matrix, whose mean absolute off-diagonal entry we prove is uniquely minimised by an orthonormal frame. The resulting regulariser, (), is four lines of PyTorch—one hyper-parameter, no parameters, no inference cost. On the canonical nine-cell benchmark it reduces cosine collapse by 99.2% and lifts the entropy-based effective rank from to of a ceiling. Against fifteen baselines spanning – it alone eliminates cosine collapse and restores near-full rank on every cell, with less collapse than the best. The separation is categorical: the two nine-cell samples do not overlap ( of pairwise comparisons, permutation ), it holds across depths and widths at an untuned , and the spectrum is left essentially flat ()—the signature of the predicted frame. Added to a standard MLPSkip block, yields the single best point of the entire benchmark on every metric, at test error for a measured training overhead. Representational collapse is thus an optimisation attractor, governable at the level of the loss.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.