Learning Dynamics Reveal Adversarial Distance and Membership Risk
Abstract
Deep neural networks memorize individual training samples to different degrees, with consequences for both adversarial robustness and privacy. Yet sample-level adversarial distance and membership vulnerability are expensive to measure at scale because they typically require iterative attacks, auxiliary models, or repeated retraining. We show that both quantities are encoded in the learning trajectory itself. We study Cumulative Sample Loss (CSL) and Cumulative Sample Gradient (CSG), which summarize how a sample's loss and input gradient evolve during training and can be collected with negligible overhead. Theoretically, we first establish a direct relation between CSL and CSG. We then derive two-sided bounds on adversarial distance from sample loss and input-gradient statistics and extend these results to trajectory-level quantities. Under stability, generalization, and model-bias assumptions, we yield a formal connection between memorization, adversarial vulnerability, and privacy leakage. These results motivate efficient sample-level vulnerability estimators from training dynamics. Our adversarial-distance proxy improves mislabeled-sample detection and achieves near-perfect duplicate detection on CIFAR-10 and CIFAR-100. For privacy auditing, we introduce CSL-Diff, which identifies samples vulnerable to LiRA, AttackR, and RMIA without auxiliary models at deployment and consistently outperforms prior artifact-based predictors. Together, our results turn learning dynamics into a reusable signal for auditing robustness, privacy, and dataset quality at the cost of standard training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.