Activation Density in Transformer Feed-Forward Networks Is Set by the Nonlinearity, Not the Width
Abstract
Two thirds of a Transformer's parameters sit in its feed-forward blocks, whose width follows from an expansion ratio fixed at 4 since 2017. Work on activation sparsity has tied that width to the fraction of the block a model uses. We find the two nearly independent. We measure the fraction of FFN units carrying the output, the activation density, with a participation ratio, which needs no threshold and so works for gated activations. Across nine Transformers trained to differ only in , density moves 7% over a fourfold range, reconciling an earlier report of sparser wide MLPs that confounded width with model scale. Density is set by the nonlinearity. Two moments of the pre-activation, with the activation function, predict it to 0.49% across non-gated nonlinearities whose densities differ by a factor of 3.6; gated activations carry a residual we quantify and leave open. The prediction's closed forms, such as for ReLU, hold only at zero pre-activation mean, which training destroys. Constraining the first-layer rows to average to zero restores that mean: ReLU density returns to within 0.65%, and the width slope vanishes, as pre-registered, for both activations whose gate is positively homogeneous. Applied to a covariance spectrum, the same participation ratio understates effective rank by a median factor of four, and by up to twenty-six, because trained networks learn a handful of massive activations. Utilization stays put while width changes, so it cannot be what sets the width; any case for a ratio must be made on capacity.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.