acceptodds
Under review as a conference paper at ICLR 2027

Piecewise Bernstein Networks: Scalable, Compatible and Universal Learnable Activations

Abstract

Learnable activation functions (like Kolmogorov-Arnold networks KAN, Polynomials, and Rational) increase model expressivity by enabling individual neurons to adapt their non-linear shapes to data. However, three critical obstacles prevent them from replacing fixed activations: a lack of scalability to large, complex architectures such as Vision Transformers, a lack of architectural compatibility, and limited expressivity in existing flexible formulations. We introduce piecewise Bernstein networks (PBN), a family of learnable activations formed by Bernstein polynomials of degree joined continuously at interior knots. We prove that PBN neurons can exactly reproduce foundational activations like ReLU, LeakyReLU, and PReLU as special cases and can approximate any other continuous activation function, rendering them into universal neurons. Against ten other activations on 13 benchmarks and in six settings, spanning function fitting, tabular MLPs, ResNets on CIFAR and ImageNet, an FT-Transformer and a vision transformer, PBN ranks first on eight benchmarks and second on the other five, and it is never behind the best fixed activation beyond seed noise. It raises ImageNet top-1 over ReLU by 1.06 points with ResNet-34 and 0.73 with ResNet-50, and the accuracy of a vision transformer by 1.0 and 2.0 points on CIFAR-10 and CIFAR-100, at 1.4 to 3.1 times ReLU’s epoch time. On every benchmark, its validation-selected configuration reaches ReLU’s final validation value in fewer epochs than ReLU.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.