Nonasymptotic P Theory for Neural Networks with Applications to Small-to-Large Generalization
Abstract
Maximal Update Parameterization (P) provides a principled framework for scaling neural networks while preserving non-degenerate feature-learning dynamics in the width limit. However, existing theory for P is largely asymptotic, relying on the infinite-width limit and leaving its behavior in finite-width regimes poorly understood. In this paper, we use the Tensor Program (TP) framework to provide a non-asymptotic characterization of gradient descent dynamics in canonical scalar-input, width- shallow settings: two-layer linear and quadratic networks, and three-layer linear networks. Our analysis provides high-probability guarantees for the finite-width approximation over each fixed training horizon. As an application, we analyze small-to-large generalization for KL-regularized domain reweighting under two-layer square-loss dynamics, showing that the learned domain weights along the width- training trajectory remain close to their infinite-width counterparts with discrepancy . Finally, synthetic and real-world experiments provide qualitative evidence of cross-width domain-weight stability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.