Counterfactual and Verbatim Memorization Disagree: From In-context Learning Theory to LLM Data Selection
Abstract
Large language models (LLMs) can memorize their training data, but whether such memorization is desirable remains an open question. Some forms of memorization are associated with verbatim training-data leakage, while others capture the contribution of individual training samples to model generalization. This creates an apparent dilemma: suppressing memorization could reduce verbatim training-data leakage, but may also hurt model utility. This paper shows that the dilemma is largely metric-dependent. We study two notions of LLM memorization: counterfactual memorization (CF-Mem), which characterizes how model behavior changes when a sample is included in or excluded from training, and verbatim memorization (VB-Mem), which measures how easily training content can be reproduced verbatim from an LLM. Through an in-context learning (ICL)-theoretic analysis, we prove that these two notions are driven by distinct mechanisms and can be inversely related: samples with higher CF-Mem tend to exhibit lower VB-Mem, while CF-Mem is positively associated with LLM generalization. This disagreement suggests that not all forms of memorization are harmful. Motivated by these findings, we propose a CF-Mem-guided LLM training data selection method. Our method efficiently estimates CF-Mem using a small proxy model and allocates more token budget to samples with higher CF-Mem scores. This yields a smaller training set that preserves samples important for generalization while reducing those associated with high VB-Mem. Across reasoning benchmarks, our method substantially reduces verbatim extraction risk and training tokens while preserving most of the full-data fine-tuning performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.