acceptodds
Under review as a conference paper at ICLR 2027

When Does Memorization Become Extraction? Frequency-Conditioned Privacy Auditing of Fine-Tuned Language Models

Abstract

Large language models are increasingly fine-tuned on sensitive data, where the same person's information may appear multiple times across records. This can cause models to memorize sensitive information, but memorization does not necessarily mean that the information can be recovered. A more serious privacy risk arises when the model begins to reproduce that information during generation. In this paper, we study how often sensitive information must appear in the fine-tuning data before it becomes recoverable. Across four language models, we vary the frequency of email addresses and phone numbers from to and compare standard fine-tuning, gradient clipping, and differentially private training. We find that the number of repetitions required before sensitive information becomes recoverable varies substantially across models. Gradient clipping can delay extraction, but it does not consistently prevent leakage at high repetition. DP-SGD at produces zero observed extraction through repetition across all four models under the attacks we test, although the utility cost varies by model. These findings show that practical privacy risk is exposure-dependent and that the transition from memorization to extraction varies across models and training mechanisms.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.