Final-Layer Jumps in Transformer-LMs: Causes, a Remedy and Interpretability Effects
Abstract
Transformer language models (LMs) exhibit a final-layer jump: hidden states change gradually across the middle layers but shift abruptly at the final layer, a pattern widely observed in existing Transformer-LMs. The jump matters for Vocabulary Projection Methods such as the Logit Lens, which read out intermediate hidden states with the final LM head. Since the LM head is trained to decode the hidden state after the jump, a large final-layer change can reduce compatibility between intermediate hidden states and the LM head, degrading readouts at all intermediate layers. Through architectural ablations, we localize the jump to the final feed-forward network and show that it can be relocated: a linear buffer layer inserted before the LM head absorbs the jump, yielding a jump-relocated counterpart. Comparing models with and without the jump, we find that intermediate-layer representations remain highly similar while Logit Lens readouts change substantially, showing that the final-layer jump alters the results of Vocabulary Projection Methods and should therefore be taken into account.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.