Distributed Bias Terms in Bias-Free Transformer Feed-Forward Networks
Abstract
Modern language models omit bias parameters from their feed-forward networks (FFNs). We show that implicit input and output biases emerge in these networks regardless, and that training organizes them by a structure that is consistent across more than twenty public dense and mixture-of-experts (MoE) models from 70M to 30B parameters. Every token's FFN input contains a large, nearly constant component along the mean input, which gives each neuron a fixed offset, and the average FFN output is added to every position; these are the input and output biases. Training arranges the neurons around the mean input: the gate vectors of most neurons point against it, while those of the few neurons that write the most are strongly aligned or anti-aligned with it. For an anti-aligned gate, the mean component drives the gate preactivation negative but does not shut the neuron off, because SiLU remains nonzero there. The output bias is therefore not the work of a few always-on neurons but a small surplus left over in almost every neuron, the largest one percent contributing between a fifth and a half. Nor are the biases constants. The channels that hold the mean input carry information about the token, including the norm of the residual stream that normalization discards, and in the middle layers context reduces how much most tokens write along the output bias. Models that do have bias parameters form the same implicit biases and keep the explicit ones at one to twenty percent of their size, and training with the mean input removed removes the structure without a measurable change in loss. In MoE models each expert has its own biases, and the output biases of different experts point in a common direction. They add up coherently in the block output, and they account for much of the similarity between experts that is measured on averaged outputs and commonly read as redundancy. Generally, our results show that mean activations are not merely nuisance statistics for mechanistic interpretability, but part of a recurring computational structure across LLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.