acceptodds
Under review as a conference paper at ICLR 2027

End-to-End Differential Privacy for Multi-stage Foundation Model Training

Abstract

Differential privacy is increasingly used in multi-stage training processes of foundation models, from initial pretraining over different fine-tuning and alignment stages, down to task-specific in-context learning. However, with current privacy accounting, we cannot report an end-to-end privacy guarantee of the final model. The two main obstacles are that (1) the different training stages may protect different privacy units, i.e., they are defined with respect to different neighborhood relations, such as fixed-length token sequences vs. task examples, or documents vs. users, and that (2) the training data overlap between the various stages might be unknown. Since differential privacy guarantees are defined with respect to their own neighborhood relation, individual per stage guarantees cannot be composed directly, and standard composition results require either perfect knowledge of data overlap or pessimistically assuming complete overlap of data between all stages. We close this gap with the first formal framework for end-to-end privacy guarantees under heterogeneous privacy units and unknown but constrained data overlap. Concretely, we formalize each stage as a private mechanism with a stage-specific neighborhood relation and derive tools that give a guarantee for the end-to-end mechanism under generic constraints on data overlap. We show that our bounds are black-box optimal in several cases, and that, surprisingly, converting to a smaller privacy unit, even for DP-SGD, can weaken the guarantee. Turning to practice, we study end-to-end accounting for Gaussian DP mechanisms, showing that our bound is easy to compute even with a large number of stages, and that the most expensive stage(s) dominate the final privacy cost. Our advances enable rigorous end-to-end privacy guarantees for modern multi-stage foundation model training pipelines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.