Anchored or Drifting: What Recursive Self-Generation Reveals About Training Data
Abstract
Large generative models are known to memorize their training data, posing severe privacy risks. Yet, current methods to detect training membership typically rely on the weak signals of a single forward pass. In this work, we find that training samples and unseen (held-out) data follow visibly different trajectories under recursive self-generation - repeatedly feeding a model's output back as its next input. Held-out samples *drift*: they lose the specifics of the original within a few steps. Training samples stay *anchored*, degrading far more slowly. We show that the membership signal this produces holds across model modalities, architectures and scales, spanning language, diffusion, and autoregressive vision models. Furthermore, these recursive trajectories provide a signal that raises membership inference TPR at FPR for nearly every attack we evaluate, roughly doubling it on the weakest baselines and still improving the strongest, which shows that a model's behavior under recursion carries membership evidence that a single query does not.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.