Rethinking Activation Functions: Reducing Bias Can Improve Learning
Abstract
Many widely used activation functions, including ReLU, GELU, Swish and Sigmoid, produce predominantly positive outputs, resulting in a nonzero post-activation mean that we term activation bias. Despite its ubiquity, the effect of activation bias on learning dynamics remains poorly understood. Using controlled models, we show that activation bias accelerates initial loss reduction but is accompanied by overshooting, non-smooth trajectories, and a steeper loss landscape. This early advantage also coincides with the model's use of task-relevant information encoded in the bias component. In contrast, reducing activation bias slows initial learning but ultimately yields a lower final loss. Experiments on larger vision and language models exhibit a similar pattern: smaller activation bias is associated with better late-stage performance. Thus, this seemingly minor property can substantially redirect learning and warrants explicit consideration in future activation function design.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.