acceptodds
Under review as a conference paper at ICLR 2027

Effective Regularity: When Point Estimates Generalize Like the Bayesian Posterior in Singular Models

Abstract

Singular learning theory explains how overparameterized models can generalize well despite having many parameters. Its main guarantees, however, concern Bayesian prediction, whereas neural networks are typically trained to a single solution using gradient descent. In singular models, these two approaches need not have the same generalization behavior. We identify a class of model-data regimes, which we call effectively regular, in which they do. Effective regularity arises when the likelihood depends on the parameters only through a lower-dimensional function class that is locally well behaved around the true function. We prove that Bayesian prediction and maximum likelihood then have the same leading generalization error, , where is the dimension of the network's function class near the truth rather than its number of parameters. We establish effective regularity for several shallow and deep linear networks and for a shallow nonlinear network. Finally, we examine empirically when gradient descent reaches the corresponding maximum-likelihood solution and therefore achieves the same rate.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.