Dense Supervision Is Not Enough: The Readout Blind Spot in Looped Language Models
Abstract
Looped language models iteratively apply the same transformer block, which enables us to vary computational depth at inference time. To ensure intermediate loops produce useful predictions, these models are usually trained with dense supervision applying a cross-entropy loss at every loop. In this work, we study a supervision failure that arises from this practice, which we term the readout blind spot. Because standard output readouts normalize the hidden state before prediction, they effectively mask the state's magnitude from the local loss gradient. Meanwhile, pre-norm residual connections continuously accumulate scale. As a result, dense supervision forces intermediate states to become predictive but entirely fails to regulate their residual norm, allowing it to grow uncontrollably across loops. Through ablations at 44M and 129M, and a matched 1.4B scaling study, we demonstrate that stabilizing looped models requires actively decoupling exit supervision from scale control. Exposing the scale to the loss (via raw readouts or norm penalties) or resetting it within the loop (via inter-loop normalization) successfully bounds the residual norm. Every form of scale control improved perplexity in our small-model recipes and at 1.4B parameters, the blind spot hold in norm readout.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.