Detectability without Interference: Efficient Canary Design for Privacy Auditing in One Run
Abstract
Privacy auditing aims to empirically assess privacy leakage in machine learning models using membership inference attacks (MIAs), and to derive lower bounds on differential privacy (DP) parameters. Standard privacy auditing inserts a sensitive sample (a canary) into many () randomized training runs, requiring the auditor to detect whether it was included in each run. In one-run auditing, many canaries are included in a single training run, avoiding the high computational cost of multi-run approaches. However, recent theory shows that canaries can interfere with one another, reducing per-canary MIA scores and thereby weakening the audit power. We propose IBIS, a framework that optimizes canaries to maximize their own score while minimizing their impact on others. IBIS combines influence-based greedy initialization with bilevel optimization, introducing a novel regularization term that encourages canaries to be orthogonal in embedding space as a proxy for reducing interference. We propose an efficient bilevel algorithm that incrementally updates a single model at each canary update, avoiding the need to retrain the model from scratch after every update. Experiments show that IBIS offers a better performance-cost trade-off than existing canary crafting approaches, improving MIA metrics while reducing computational cost by up to two orders of magnitude.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.