Localizing Verbatim Memorized Spans in Language Model Outputs
Abstract
Large language models (LLMs) can reproduce training data verbatim, raising copyright and privacy concerns. To detect such reproduction, existing methods typically frame the problem as membership inference, classifying entire sequences as members or non-members. However, a long generated sequence may contain only a short memorized passage, for example, roughly 70 verbatim tokens within an up to 512-token output. Sequence-level aggregation can therefore dilute the memorization signal. While token-level scoring methods exist, they commonly flag tokens independently. Our key insight is that verbatim memorization should be treated as a structured span-localization problem rather than a sequence-level or independent-token detection problem. Unlike membership inference, which asks whether a candidate text was used in training, our task asks where verbatim matches occur within generated text. We make three contributions. First, we develop a benchmark for predicting the boundaries of verbatim spans and demonstrate the benefit of contiguity-aware extraction: on Hubble models, applying it to InfoRMIA improves mean intersection over union (IoU) from 0.326 to 0.683. Second, we introduce Token Span Localization (TSL), which converts token-level scores into contiguous span predictions using two scoring functions, Final Loss (FL) and Sample Loss Difference (SLD). Third, we provide a theoretical analysis showing that these scores offer complementary evidence: FL captures how likely the final model is to reproduce a complete span, while SLD measures how much its likelihood has increased since a reference checkpoint. On GPT-2 XL, TSL achieves 0.851 mean IoU (with SLD) and 0.984 detection AUC (with FL), a gain of 0.122 over the strongest sequence-level baseline (0.862). Across eight Hubble configurations, TSL obtains 0.986-0.993 AUC, an average gain of 0.207 over InfoRMIA (0.672-0.850), compared with 0.601–0.796 for the strongest conventional sequence-level baseline in each configuration. TSL also achieves the highest detection AUC among the evaluated methods on 12 of 14 open-weight models tested against an external corpus. Together, these results establish TSL as a more accurate and fine-grained approach to auditing verbatim memorization in LLM outputs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.