A Structural Approach to Path-Norm-Based Generalization Bounds for Neural Networks
Abstract
While the path-norm has emerged as a theoretically grounded measure of neural network generalization, existing generalization bounds often rely on restrictive assumptions—such as the positive homogeneity of activation functions—and predominantly target basic feedforward or residual architectures. In this paper, we develop an analytical framework that broadens the scope of path-norm theory to modern neural network architectures without relying on potentially loose covering-number estimates. First, by introducing the Absolute Rademacher Complexity, we remove the traditional positive homogeneity requirement and extend path-norm-based bounds to arbitrary Lipschitz-continuous activation functions. Second, we establish continuous-limit path-norm bounds for Neural Ordinary Differential Equations (Neural ODEs). By proving the uniform convergence of Rademacher complexities under Euler discretization, we characterize their generalization complexity through a differential comparison system. Finally, we decompose the self-attention mechanism and combine an element-wise matrix inequality with vector contraction to derive explicit, layer-wise bounds for Transformers. Our structural decomposition approach avoids the theoretical bottlenecks of traditional covering-number methods and provides architectural insights into how continuous-depth dynamics and attention operators contribute to model capacity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.