Sem2Pix: Privacy Leakage from Gradients via Semantic-to-Pixel Fine-Grained Guidance Diffusion Models
Abstract
Federated learning (FL) enables collaborative model training among multiple clients without requiring them to expose their raw images. Several studies have introduced attacks known as gradient inversion attacks (GIAs), which exploit the gradients shared during training to reconstruct clients' raw images. Recent work has shown that training-free diffusion models can transform Gaussian noise into recovered private images under guidance from shared gradients without requiring knowledge of the clients' data distribution. Our observations suggest that (i) guidance should constrain semantic information, such as global shape and color distribution, and (ii) additional guidance is needed to recover fine-grained details, which are particularly important in sensitive domains such as healthcare. In this paper, we present Semantic-to-Pixel Fine-Grained Guidence (Sem2Pix), which leverages vision–language models (VLMs) to enhance guidance for diffusion-based GIAs. First, we introduce VLM guidance for reverse diffusion to constrain the reconstructed image to be semantically meaningful. Second, we enhance the recovery of fine-grained details through stage-wise gradient guidance, which dynamically matches the layer-wise gradients induced by the predicted clean image with the corresponding shared client gradients. With these two steps, Sem2Pix can recover high-resolution images with high fidelity in practical federated learning scenarios involving defense mechanisms, larger batch sizes, and multiple local updates.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.