Magnitude-Induced Rank Collapse in FFN Value-Vector Gradients
Abstract
Each column of the FFN down-projection maps an intermediate channel to the residual stream in a Transformer model. How individual training sequences contribute to updating these columns is obscured by minibatch aggregation and model- or layer-level gradient summaries. We therefore characterize each column's per-sequence gradient matrix using normalized energy effective rank (NEER), which responds to both directional correlation and gradient-norm imbalance. We first observe that columns within the same layer exhibit a pronounced low-rank tail across optimization, showing that sequence contributions can differ sharply across parameter vectors. We further find that low NEER often reflects gradient-magnitude imbalance rather than directional alignment, with many low-NEER columns retaining substantial directional diversity. To trace the source of these outliers, we use the standard activation–backpropagation factorization. It reveals that large column-specific activations and strong residual-space backpropagated signals coincide at the same token positions, whereas neither factor alone explains the concentration. Together with a simplified optimization analysis, these findings suggest that high-magnitude sequence gradients in some columns can suppress otherwise diverse directions and limit their contribution to optimization. Motivated by this analysis, we test parameter-local selective clipping of dominant sequence gradients in low-NEER columns. Across Llama-3.1-8B, Llama-3.2-3B, and Mistral-7B, selectively rebalancing dominant local contributions yields consistent downstream gains, suggesting that reducing parameter-local gradient imbalance can be beneficial during fine-tuning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.