Activation-Aware Initialization for Odd-Sigmoid Networks
Abstract
Activation functions and weight initialization are tightly coupled, especially in deep networks with bounded sigmoidal nonlinearities. Standard zero-mean Gaussian initializations, including Xavier, He, and edge-of-chaos variants, often become brittle in deep and narrow fully connected networks, where forward activations collapse and backpropagated gradients decay. We study this problem for a class of bounded, odd, monotone activations whose slopes decrease away from the origin, which we call odd-sigmoid activations. For an activation , our initialization sets a diagonal gain using the critical slope and adds Gaussian noise whose scale is calibrated by a terminal sign-flip surrogate. The scalar analysis gives a closed-form relation between the surrogate sign-flip target, depth, and noise scale, while an empirical depth schedule transfers this calibration to finite-width networks. Across deep fully connected networks on image benchmarks and PINN benchmarks, the proposed diagonal-plus-noise initialization is substantially less sensitive to depth, width, activation scaling, and learning-rate choice than Gaussian i.i.d. baselines. The results suggest that sign-flip calibration is a practical tool for training deep odd-sigmoid networks, while also highlighting the need for structured initialization beyond variance-only Gaussian schemes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.