acceptodds
Under review as a conference paper at ICLR 2027

M-Drift: Membership Inference via Gradient-Induced Mechanistic Drift

Abstract

Vision-language models (VLMs) and diffusion models are often trained on large web-scale datasets, making it difficult to determine whether a particular example was used during training. We introduce M-Drift, a white-box membership inference framework that uses a model's response to a small, reversible gradient-ascent perturbation as membership evidence. For each candidate example, M-Drift applies a single loss-increasing nudge and measures the resulting changes in loss, internal representations, and cross-modal interactions. For VLMs, we capture representation and cross-attention drift; for diffusion models, we aggregate denoising drift across sampled timesteps and noise realizations. We further use sparse feature projections to localize the internal directions most affected by the intervention. Across multiple VLM and diffusion backbones, M-Drift consistently improves membership detection in low-false-positive regimes over strong output- and gradient-based baselines. These results suggest that intervention-induced representation drift provides a useful and interpretable signal for auditing memorization in modern generative models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.