Boosting Diversity and Relevance: Privacy-Preserving Demonstration Synthesis for In-Context Learning
Abstract
Large language models (LLMs) rely on contextual information embedded in demonstrations to perform effective in-context learning (ICL), yet using demonstrations drawn from a private data pool may expose sensitive information when interacting with LLM servers. Existing demonstration synthesis methods based on differential privacy (DP) can mitigate such privacy risks, but they struggle to balance privacy and utility. Particularly, these methods often generate demonstrations with limited diversity and task relevance, while DP noise injected into token sampling distributions introduces semantic errors that accumulate during autoregressive generation. Motivated by these limitations, we propose DPD, a novel privacy-preserving algorithm to generate high-utility demonstrations under DP guarantees. To address diversity and relevance gaps, we develop a task-aligned encoder that maps private demonstrations into a task-aware embedding space, where a DP coverage-aware strategy privately releases coverage-weighted centers and retrieves candidate demonstrations for promoting semantic coverage. To mitigate error accumulation, we avoid perturbing token autoregressive generation and instead introduce the Joint Exponential Mechanism to efficiently extract DP-based keywords from candidates, guiding the advanced LLM to synthesize demonstrations with high utility. The overall DPD framework provides demonstration-level DP. Extensive experiments on standard benchmarks show that DPD outperforms token-level DP demonstration synthesis baselines while effectively mitigating membership inference attacks, providing an effective solution for privacy-preserving ICL.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.