acceptodds
Under review as a conference paper at ICLR 2027

Target Scale Is Readout Scale: One Ratio Orders How Long Collapse Lasts

Abstract

Many self-supervised methods train an encoder through a linear readout that regresses a dense target. When the encoder ends in a layer normalization, the readout's output at initialization has a norm set by the widths alone, regardless of the target. Scaling the target is therefore equivalent to scaling the readout, and what matters is the ratio between the readout's initial output and the target's norm. The standard recipe, a unit-gain layer normalization with a fan-in initialized readout and a unit-variance target, fixes at every width. We test what happens when it does not, in a controlled case study: MelHuBERT trained on LibriSpeech with regression targets of different scales. Rank collapsed early in every run with a default readout, and ordered how long the collapse lasted: about a thousand steps at and more than twelve thousand at . Of 94 pairs of runs at different ratios, 90 are ordered as predicts, none against it, and 4 are undetermined because neither run had recovered. Most of the collapse is a shared offset, with 98% of the representation's energy in its mean over frames; with the mean removed rank falls by over a third, and further under a softmax readout, which builds no offset. At 8,500 steps orders both the offset and the rank that remains without it. In 500-step runs at , a sixty-fold change in the readout's learning rate moved rank by two, while lowering restored 100 to 419 dimensions. At large the gradient at the encoder's output is dominated by a term that does not contain the target, and an approximate closed form for its direction matches simulation to 0.3%. MAE (0.89) and data2vec (0.58 to 1.63) sit close to . Distance from initialization measured by CKA, a common diagnostic, can miss the collapse: CKA centers the representation, and two runs 0.12% apart on it differ sevenfold in rank, twofold with the mean removed. Checking costs one batch, before training starts.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.