Guided Diffusion for Trustworthy Synthetic Text Generation
Abstract
The rapid evolution of Large Language Models is fueled by massive datasets, yet this reliance poses severe privacy risks, as models can inadvertently memorize and even directly output sensitive training data. Synthetic text generation offers a promising solution to this challenge; however, existing methods struggle to reconcile rigorous privacy guarantees with high data utility. In this paper, we propose a new paradigm for private synthetic text generation: Guided Diffusion. Leveraging the unique steerability of Diffusion Language Models, we introduce novel sampling and re-masking procedures that actively guide the iterative denoising steps. These mechanisms steer the generative process to reconstruct fluent text that adheres to a target semantic distribution. We demonstrate that our framework is naturally compatible with Differential Privacy (DP): once the semantic embeddings of the private corpus are protected, the DP guarantee carries over to the generated text by post-processing. Empirically, we show that our approach generates high-quality synthetic text with precise semantic alignment. Furthermore, we validate its effectiveness in DP synthetic text generation: across two datasets and two diffusion backbones, it outperforms Private Evolution and DP-SGD-based fine-tuning in distributional fidelity and downstream utility, while generating significantly faster than Private Evolution.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.