Statistical guarantees for robust deep learning under heavy-tailed noise and covariates
Abstract
Robust loss functions are widely used in deep-learning pipelines to reduce sensitivity to outliers, heavy tails, and other forms of data corruption. In this paper, we support these empirical findings by providing statistical guarantees. We derive oracle inequalities for empirical risk minimization with Lipschitz losses over nonconvex function classes, including feedforward neural networks. Our results only require weak moment assumptions on the covariates, while remaining noise-agnostic. We also study non-Lipschitz target losses through training with Lipschitz surrogates. The resulting bounds provide guarantees for the original target risk and make explicit the estimation–approximation tradeoff, in accordance with the classical statistical robustness–efficiency perspective. The framework recovers Huber-type procedures and extends to more general losses. Numerical experiments under heavy-tailed corruption illustrate the benefits of Lipschitz surrogate training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.