Overfitting Management in Deep Networks: A Sensitivity Correction to the Training Loss
Abstract
Overfitting in deep neural networks (DNNs) is handled by a combination of techniques: an L2 penalty and dropout, each with its own hyperparameter, and early stopping on a held-out validation set. All are set from outside the fit: tuned or monitored on held-out data rather than estimated within it. In classical statistics, by contrast, the hyperparameter controlling L2 can be learned from the data. Empirical Bayes estimates it as a variance component while fitting the model, and the penalized objective is minimized to convergence with no stopping rule. This paper pursues the same goal in a DNN: estimating from the training data, within SGD, the strength at which a fit neither overfits nor underfits. The classical construction does not transfer unchanged, and the adaptation and extension it requires are the substance of the work. The difficulty is that the very expressive power that makes a DNN useful enables it to hide overfitting. A DNN can satisfy an L2 penalty by holding a weight small while amplifying the gain applied downstream, so influence grows while magnitude does not, and the penalty stops binding. The same freedom opens the door to a kind of overfitting that neither the penalty nor its classical estimate can see, and that is why in DNNs L2 is left as a hyperparameter and monitoring is needed to decide when to stop. We add a sensitivity correction to the L2 penalty, pricing what that penalty cannot see: the sensitivity the network grants itself. The adjusted loss has its roots in the hierarchical likelihood of Lee and Nelder (in a linear mixed model it reduces exactly to REML, so the correction is an extension of REML to DNNs). Regularization strength is then estimated during ordinary training, and no validation set is needed to set it or to stop training. The quantity controlled is the model’s effective degrees of freedom, the capacity a fit uses rather than the parameters it owns. We call the approach OMID, for Overfitting Management in Deep Networks, and focus first on embedding-dominated architectures. We test OMID in simulation, where the data-generating process is known, and on the public Criteo and Avazu click-through-rate benchmarks. Both held-out loss and ranking accuracy improve. Under tuned L2 with early stopping, held-out loss reaches its best within a few epochs and rises sharply thereafter, and that best is not the best available. With the correction, loss stops turning: it keeps falling as training continues and ends below that value at every data size tested, with the largest gains where data is scarcest. Training and held-out performance no longer diverge, and further passes bring diminishing returns rather than damage. Expected calibration error drops by factors of three to twenty. We close with the extension now under way, to sequence models and large language models, where transformer and self-attention architectures change how sensitivity must be addressed.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.