M-Drift: Membership Inference via Gradient-Induced Mechanistic Drift
Abstract
Vision-language models (VLMs) and diffusion models are often trained on large web-scale datasets, making it difficult to determine whether a particular example was used during training. We introduce M-Drift, a white-box membership inference framework that uses a model's response to a small, reversible gradient-ascent perturbation as membership evidence. For each candidate example, M-Drift applies a single loss-increasing nudge and measures the resulting changes in loss, internal representations, and cross-modal interactions. For VLMs, we capture representation and cross-attention drift; for diffusion models, we aggregate denoising drift across sampled timesteps and noise realizations. We further use sparse feature projections to localize the internal directions most affected by the intervention. Across multiple VLM and diffusion backbones, M-Drift consistently improves membership detection in low-false-positive regimes over strong output- and gradient-based baselines. These results suggest that intervention-induced representation drift provides a useful and interpretable signal for auditing memorization in modern generative models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.