Noise-free Provable Trustworthy Machine Learning via Distribution-Specific Indistinguishability
Abstract
Indistinguishability guarantees in trustworthy machine learning are usually obtained by injecting noise, most prominently through Differential Privacy (DP). We develop a noiseless alternative for data-influence problems such as poisoning, memorization, and copyright protection. Let be the part of the training data whose influence we want to control (poisoned samples, a sensitive record, or a protected artwork), and let be the rest. A model trained on alone never sees and is taken as a safe reference. Our goal is to make the model trained on statistically indistinguishable from this reference, so that the reference's trustworthiness carries over. Operationally, instead of injecting algorithmic randomness such as noise, we exploit input-wise randomness: the benign data is modeled as a random sample from an unknown distribution, while may be chosen adversarially after observing . We formalize this as Distribution-Specific Indistinguishability (DiSI): a DiSI algorithm then transfers the reference's high-probability trustworthiness to the model trained on . We instantiate DiSI through consistency-robust gradient operators that make the target and reference updates agree exactly, except on a failure event whose probability decays exponentially in the number of sources. Across image and text backdoor benchmarks, DiSI sharply reduces attack success while preserving far more utility than noise-based baselines; in generative-model fine-tuning, it also reduces outlier memorization, membership leakage, and resemblance to protected artworks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.