Same Privacy Guarantee, Different Robustness: How Gradient Clipping Shapes Adversarial Training
Abstract
Trustworthy machine learning requires addressing multiple threats, including privacy leakage and adversarial perturbations. Differential privacy and adversarial robustness provide complementary forms of protection against these risks. Combining techniques for these goals within the same training procedure can change how a model learns. In private gradient-based training, bounding each example's contribution calibrates privacy noise but can also alter the balance between classification and robustness. The loss coefficients and privacy budget alone therefore do not specify the update used to learn these goals. We compare two ways of limiting gradient contributions: bounding classification and robustness gradients jointly or separately, with the same objective and privacy calibration. We analyze the bounded updates before noise is added, then use controlled fine-tuning experiments in two image-classification settings to test their robustness consequences under hard and smooth clipping. Our analysis shows that preserving the balance between loss components within each example need not preserve the combined update across examples. Across both settings and clipping maps, matching magnitudes to joint bounding while retaining componentwise directions improves held-out AutoAttack accuracy; restoring directions alone is less consistent. Joint exceeds Componentwise by over three percentage points in one setting, but their ordering changes with the setting and attack radius. Contribution magnitudes affect learned robustness even when per-example gradient directions are preserved.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.