acceptodds
Under review as a conference paper at ICLR 2027

Optimization Risk Bounds for Kolmogorov–Arnold Networks Trained by DP-SGD with Correlated Noise

Abstract

The theoretical understanding of differentially private stochastic gradient descent (DP-SGD) with temporally correlated noise remains limited, particularly for non-convex neural network training. As a first step, we study two-layer Kolmogorov–Arnold Networks (KANs), a recently introduced architecture with learnable spline-based edge functions. We establish the first optimization risk bounds for clipped mini-batch DP-SGD with correlated noise in this setting, with explicit dependence on temporal correlation, clipping, mini-batch sampling, and network width. Existing arguments fail for three reasons: temporal dependence breaks the conditional-centering step; projection obstructs the cross-iteration cancellation of correlated perturbations; and active clipping breaks the empirical-gradient structure. We address these issues through shifted and auxiliary dynamics, a weighted empirical loss, and a high-probability localization argument. Our bound shows that temporal correlation reduces the leading noise terms, the clipping threshold enters essentially through an effective step size, and private training admits an explicit network width range. Experiments on synthetic data and MNIST support the predicted optimization effect of temporal correlation. As an application, we derive population risk guarantees via algorithmic stability. Our framework recovers non-private mini-batch SGD, independent-noise DP-SGD, and their full-batch counterparts as special cases.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.