Improving the Empirical Risk Bound for Private Learning under Interpolation
Abstract
Differentially Private Gradient Descent (DP-GD) provides a principled approach to training machine learning models while preserving data privacy by injecting Gaussian noise into model updates and clipping sample-wise gradients. Existing literature has established joint theoretical guarantees for differential privacy and excess empirical risk under classical assumptions such as (strong) convexity and the Polyak-Łojasiewicz (PL) condition, with excess empirical risk bounded by or worse. In this paper, we study DP-GD in the *interpolation regime*, a key setting in modern machine learning where the training data can be fit exactly. We show that interpolation enables substantially stronger optimization guarantees, transforming polynomial-type empirical-risk bounds into exponential-type bounds in and , while still preserving differential privacy. Specifically, leveraging a property of interpolation — *automatic gradients reduction*, we propose to synchronously decay the standard deviation of the injected Gaussian noise and the gradient clipping threshold, at an appropriate exponential rate. First, in the classical settings considered in prior work, we show that interpolation improves the empirical risk to . We further extend our analysis to non-convex neural networks and establish the same exponential empirical risk bound, which is also preserved under stochastic training. Finally, empirical evaluations verify the effectiveness of the proposed hyperparameter decay mechanism and demonstrate improved convergence compared with standard DP-GD.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.