acceptodds
Under review as a conference paper at ICLR 2027

Extractable Memorization From First Principles

Abstract

Recent work on extractable memorization in language models suffers from two contrasting validity problems. Some studies overstate extraction, for example, by using sequences too short to distinguish memorization from predictability. Others imply that extraction is unreliable evidence because models can reproduce real-world text they weren't explicitly trained on. Both overlook what makes a valid extraction claim: the model must generate a training sequence with high enough probability to indicate memorization. To determine what's high enough, one has to perform a matched comparison, measuring generation probabilities of training and comparable non-training sequences. Since non-training sequences can't have been memorized, their probabilities are a baseline for predictability; exceeding this baseline is evidence of memorization. We formalize matched comparisons with (1) a conformal test calibrated to a chosen false-positive rate when sequences are sampled from populations, and (2) a single-document census that calibrates against a matched non-training document. Matched comparisons enable rigorous, calibrated memorization claims and clarify where prior setups have validity issues. On Wikipedia, OLMo 2 32B reproduces non-training 10-token suffixes roughly 24% as often as training ones: that share reflects false positives, not memorization. For Llama 3.1 70B on books, calibrated census thresholds reach as low as , supporting memorization claims for sequences no feasible sampling budget would extract. We therefore refine "extractable memorization" to require both a valid memorization claim and near-certain generation within a realistic budget.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.