acceptodds
Under review as a conference paper at ICLR 2027

Reconstruction of Private Speech using Generative-Prior-Guided Gradient Inversion

Abstract

Federated learning enables collaborative training without centralising raw audio, yet the gradients shared during training can leak sensitive information, including speech content and speaker identity. Existing audio gradient inversion attacks primarily reconstruct inputs directly in the feature space and have been studied mostly under relatively favourable inversion settings. In this work, we introduce a generative-prior gradient inversion framework for audio, which constrains reconstruction using a learnt mel-spectrogram prior and employs a neural vocoder for waveform recovery. To prevent prior memorisation of victim identities, the generator is trained using an identity-disjoint speaker split. We evaluate the attack on Speech Commands V2 and AudioMNIST under a range of increasingly challenging settings, including gradient clipping with Gaussian perturbation, compressed MFCC features, trained victim models, and different network architectures. Our results show that the main advantage of the generative prior emerges as the inversion problem becomes less informative. Under the strongest tested gradient perturbation, speaker recovery remains at 84.0% on AudioMNIST and 96.5% on Speech Commands, compared with 18.0% and 29.3% for the GAN-free baseline. These findings demonstrate that degraded speech reconstruction does not ensure speaker privacy, as identity leakage can persist through generative priors.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.