Self-Consistent and Extremal Activation Functions: A Variational Theory of Activation Design
Abstract
Which activation function should a trained network have? We pose activation design as a variational problem on the affine Sobolev space –minimize empirical risk plus an penalty over , with weights held fixed–and characterize the solutions structurally. Minimizers exist, and every critical point satisfies the self-consistent representer equation , with the Green's function of and the network's data-averaged perceptual measure: the optimal activation is a exponential spline whose knots are the pre-activation values the network itself produces under –not parameters, but a fixed point. Above an explicit threshold, the problem is strongly convex on a ball containing every minimizer; the fixed point there is unique, and Picard iteration gradient descent at step converges geometrically, at one backward pass per sample per iterate. An extremal strand explains piecewise-linearity: among all activations leaving every prediction of a trained network unchanged, the piecewise-linear interpolant minimizes the uniform first-order attack-surface bound, and under gradient fidelity the extremal has bang-bang curvature — a perfect spline in the sense of Favard and Karlin. Finally, parametric activation families (adaptive splines, Kolmogorov–Arnold edges) are Galerkin sections that -converge to the full problem at the rate . Every prediction is verified on trained MNIST networks over five seeds: the contraction exponent within of theory, knot self-consistency at grid resolution, the rate, and machine-precision behavioral equivalence of the piecewise-linear representative.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.