Data-Dependent Complexity Bounds for Saturating Networks via the Pushforward Density
Abstract
We derive a new Rademacher complexity bound for neural networks with saturating activations that requires neither input truncation nor activation surrogates. Our key insight is that the data distribution regularizes the inverse Jacobian’s boundary singularity: under mild tail conditions, the pushforward density remains bounded despite divergent inverse derivatives. We characterize this boundedness via asymptotic analysis and show that depth increases the density supremum only polynomially without worsening the singularity, extending to high-dimensional and dimension-reducing maps. Experiments confirm that our density-based measure correlates more closely with generalization gaps and yields strictly tighter bounds than prior methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.