Power-Law Data Spectra as Implicit Regularization in Learning
Abstract
Neural networks often exhibit substantially different learning performance on real-world and synthetic data, suggesting that structural properties of real data may play an important role in learning. Large-scale datasets used in modern neural-network applications, including images and language, commonly exhibit approximately power-law covariance spectra. Such spectral structure may therefore contribute to the learning properties observed on real data. However, how this ubiquitous power-law structure affects learning remains unclear. Here, we show that power-law spectral structure can act as an implicit regularizer of learning. To characterize this effect quantitatively, we analyze double-descent learning curves in ridge regression while systematically varying the power-law exponent. We establish an explicit relation between the power-law exponent and the ridge penalty that reproduces the leading behavior of the double-descent peak in the weak-ridge regime. This relation shows that increasing the power-law exponent suppresses noise amplification in a manner analogous to increasing explicit regularization. Furthermore, we find that the spectral structure affects the dependence of prediction error on sample size beyond the interpolation regime, reshaping the overall learning curve. Establishing the power-law structure of real-world data as an intrinsic source of implicit regularization that shapes learning performance, our findings advance the theoretical understanding of the role of data structure in learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.