Self-Evolving Private Evolution for Differentially Private Synthetic Text Generation
Abstract
Differentially private synthetic text generation enables data sharing and text mining applications while providing formal privacy guarantees. Recent advances leverage large language models (LLMs) as strong generative priors for producing fluent and semantically rich synthetic text. Among these methods, Private Evolution (PE) is a promising approach that incorporates LLM inference APIs to generate and refine synthetic samples without expensive fine-tuning. PE iteratively improves synthetic data by generating variations, evaluating candidates through private voting, and selecting high-quality samples for the next iteration. However, existing PE methods discard these privatized evaluation signals after each iteration, despite their potential utility for future rounds. We propose SelfEvo-PE, which reuses available privatized signals as supervision to train lightweight evaluators that guide future evolution steps without additional privacy cost. This mechanism allows SelfEvo-PE to perform more evolution steps under the same privacy budget. Our experiments on three datasets show that SelfEvo-PE outperforms the conventional PE (Aug-PE) under limited-data and strict-privacy settings, improving FID and MAUVE by up to 15% and 20%, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.