Reconstruction of Private Speech using Generative-Prior-Guided Gradient Inversion
Abstract
Federated learning enables collaborative training without centralising raw audio, yet the gradients shared during training can leak sensitive information, including speech content and speaker identity. Existing audio gradient inversion attacks primarily reconstruct inputs directly in the feature space and have been studied mostly under relatively favourable inversion settings. In this work, we introduce a generative-prior gradient inversion framework for audio, which constrains reconstruction using a learnt mel-spectrogram prior and employs a neural vocoder for waveform recovery. To prevent prior memorisation of victim identities, the generator is trained using an identity-disjoint speaker split. We evaluate the attack on Speech Commands V2 and AudioMNIST under a range of increasingly challenging settings, including gradient clipping with Gaussian perturbation, compressed MFCC features, trained victim models, and different network architectures. Our results show that the main advantage of the generative prior emerges as the inversion problem becomes less informative. Under the strongest tested gradient perturbation, speaker recovery remains at 84.0% on AudioMNIST and 96.5% on Speech Commands, compared with 18.0% and 29.3% for the GAN-free baseline. These findings demonstrate that degraded speech reconstruction does not ensure speaker privacy, as identity leakage can persist through generative priors.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.