acceptodds
Under review as a conference paper at ICLR 2027

Max-Total-Flow Generalization Bounds for DAG-Structured ReLU Deep Neural Networks

Abstract

Understanding generalization in over-parameterized neural networks remains a fundamental challenge. While recent path-norm frameworks address general directed acyclic graph (DAG) networks with skip connections, branching, and merging, they require computing astronomically large path-norms that render generalization bounds vacuous in practice. In this study, a new perspective is proposed by interpreting Karush–Kuhn–Tucker (KKT) optimality conditions as a flow conservation law over the DAG. At a KKT point, the squared Frobenius norms of weight matrices define non-negative flows on edges that satisfy classical conservation equations at every interior node. Using flow decomposition, we prove that at every KKT point the total parameter norm is bounded in terms of the longest source-to-sink path length of the DAG and the regularization parameter . Combining this norm control with a held-out validation set and a change-of-measure argument under a Gaussian perturbation of the weights with noise level , we show that, with probability at least , the validation loss of the trained network exceeds its expected training loss by at most , for training and validation samples. The results cover binary and multi-class classification and regression, and extend to Leaky ReLU activations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.