ICL-PrivEval: A Data-Dependent Framework for Privacy Leakage in In-Context Learning
Abstract
While in-context Learning (ICL) uses exemplars in the prompt to adapt pre-trained Large Language Models (LLMs) to downstream tasks without fine-tuning, these exemplars can carry sensitive information and be potentially leaked from the LLM's output. While prior works have explored this leakage through Membership-Inference Attacks (MIAs), there lacks a systematic evaluation of the factors that contribute to unintended privacy leakage from ICL. To this end, we propose ICL-PrivEval, an evaluation workflow for measuring this leakage. At its foundation is ICLInf, an exemplar-level leakage metric inspired by data-dependent differential privacy and counterfactual influence, which captures how the LLM's answer distribution shifts when each exemplar is removed. Using ICL-PrivEval, we study factors that affect ICL leakage and apply ICLInf as a privacy auditor, yielding an effective audit of sampling-based DP-ICL methods up to ε = 10.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.