Where Membership Lives: Task Abstraction, Exemplar Retrieval, and Privacy Leakage in In-Context Learning
Abstract
In-context learning (ICL) creates a different privacy problem than pretrained and fine-tuned language models: adaptation occurs without parameter updates, while sensitive demonstrations remain explicitly available in the active context and can be re-accessed during inference. We study this contextual membership and ask where membership information resides, how it differs from parameter-updated membership, and how it can be inferred under different model accesses. We introduce controlled behavioral interventions that separate target-specific exemplar information, surface-token sensitivity, and shared task abstraction. A matched ICL–Full-SFT comparison experiment further reveals different privacy footprints: under paraphrased target-associated exposure, ICL yields stronger pooled member- ship signal than SFT (0.869 versus 0.688) and a differently concentrated token-level signal. Finally, replacing private demonstrations with differentially private synthetic demonstrations reduces the evaluated attack to approximately chance, while increasing prediction loss. Together, these results identify active-context exposure as a distinct source of privacy leakage—different from the persistent parameter traces studied in pretraining and fine-tuning—and motivate privacy analyses and defenses specifically designed for ICL.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.