acceptodds
Under review as a conference paper at ICLR 2027

Measuring and Mitigating Private Data Leakage in Automated Prompt Optimization

Abstract

Automated prompt optimization (APO) has emerged as a lightweight method for adapting a large language model (LLM) to perform well on a particular task. APO edits and refines input prompt text, a sample-efficient and flexible alternative to parameter-based fine-tuning. For tasks that involve interacting with users' private data, further prompt quality improvements are possible by performing APO directly on real users' examples rather than public or synthetic data. However, this exposes the risk of private information leaking into the optimized prompt. In this paper, we demonstrate that private data leakage can indeed occur when APO is run using training data containing private information. We introduce empirical privacy auditing methods to quantify this leakage based on carefully calibrated matching of token windows between training data examples and the optimized prompt. We then present some simple modifications to APO, specifically the meta-prompts used by the mutator and autorater models, that significantly reduce private data leakage without hampering the quality of the optimized prompt. Comparison of different leakage mitigation strategies in terms of their privacy-utility trade-offs shows that private data leakage through APO can be mitigated with minimal impact on the performance of the optimized prompt.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.