acceptodds
Under review as a conference paper at ICLR 2027

Black-Box Membership Inference Attack for Vision-Language Models Using Self-Memory Probing

Abstract

The rapid advancement of vision-language models (VLMs) increasingly depends on large-scale visual corpora, raising concerns about privacy risks and copyright compliance. Membership inference attacks (MIAs) provide a means to audit whether individual samples were used for training. However, existing VLM MIAs either require white-box or gray-box access to logits and token probabilities or rely on output-only proxies derived from task-specific visual interventions or surrogate models. These access and design dependencies limit their applicability to commercial VLMs deployed under the black-box Generation-as-a-Service (GaaS) paradigm, which exposes only final outputs. To address this gap, we propose Self-Memory Probing (SMP), an output-only black-box framework grounded in the self-memory effect. This effect manifests as greater consistency and concentration in responses to member images than in those to non-members. SMP probes this distinction by eliciting multiple outputs with semantically equivalent prompts to construct candidate-conditioned and non-member reference distributions. It then casts membership inference as a hypothesis test and uses Maximum Mean Discrepancy (MMD) to quantify the discrepancy between these distributions. Consequently, SMP derives membership evidence directly from the model’s native responses, without image perturbations, surrogate models, or token-level statistics.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.