Quantifying Canary Score Dependence in One-Run Unlearning Audits
Abstract
One-run privacy audits for machine unlearning embed canaries in a forget set, apply an unlearning operator, and aggregate per-canary membership scores into a confidence statement, assuming the scores contribute independent information about the run. We show this assumption fails in a structured way. Adapting the intraclass-correlation (ICC) decomposition from biostatistics, the marginal ICC(1,1) is near zero for conservative operators (NPO, RMU) while the conditional ICC(3,1), on canary-centered residuals, is substantial—– for conservative operators, – for gradient ascent, all converged: a marginal diagnostic is blind to the dependence that matters for per-canary procedures. Within the dependence that survives canary centering, the run-level share dominates, with a smaller entity-block component (permutation ); the mechanism is not identified by this statistic. At the conditional design effect reaches for converged gradient ascent—400 canaries carry only independent observations (scalar) or (eigenvalue bound)—and independence-assuming mean-MIA intervals under-cover (marginal tier, convention), while the per-canary-count statistic mildly under-covers at – naive but is over-conservative () once corrected; for conservative operators a fresh canary draw expects near-zero transferable dependence (ICC(2,1): NPO , RMU ). At the highest coupling strength (converged GA), the correctly scaled Level-1 correction reaches full coverage on this cell (, ; unchanged under residual-run resampling), so the auditor proceeds rather than declining. The framework requires – calibration runs for both detection and magnitude estimation, so it is a *diagnostic* that tells an auditor when the one-run regime is valid.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.