Interpretable Differentially Private Continual Learning with Prototype-Based Classifiers
Abstract
Continual learning (CL) under differential privacy (DP) involves two competing objectives: a model must learn from new tasks while avoiding catastrophic forgetting, and the training pipeline must preserve privacy. Requiring the model’s decisions to be interpretable introduces an additional competing objective. These trade-offs become particularly pronounced when no sensitive state, such as memory buffers, can be retained across tasks. We introduce a DP prototype-based CL framework built on pretrained backbones, where each class is represented by a single DP prototype and predictions are made by nearest-prototype classification in a shared embedding space. This eliminates the need to store samples while making the decision rule inherently interpretable. We control the CL trade-offs through novel DP-friendly loss functions adapted from prior CL work. On the Split-CIFAR-100 and Split-ImageNet-R benchmarks, our method achieves similar or higher accuracy than a state-of-the-art interpretable DP CL baseline while remaining competitive with a stronger but non-interpretable DP CL baseline across challenging CL settings. Finally, we introduce two complementary methods for interpreting the learned DP prototypes: one based on publicly available pretraining data and another using synthetic samples generated by a generative model.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.