acceptodds
Under review as a conference paper at ICLR 2027

COMPASS: Membership Inference Attack Against Vision Large Language Models via Adaptive Noise Calibration

Abstract

Vision large language models (VLLMs) are increasingly deployed in privacy-sensitive domains, raising concerns about the exposure of their training data. The strongest existing gray-box membership inference attacks (MIAs) against VLLMs commonly rely on entropy statistics or generated descriptions, requiring costly response generation while providing limited membership separation. In contrast, perturbation-based methods measure how model responses change when images are perturbed, but have mainly focused on large perturbations, leaving VLLM membership behavior under smaller perturbations poorly understood. In this work, we systematically characterize VLLM membership behavior under controlled Gaussian perturbations across the full range of noise levels at fine resolution. Our analysis reveals that both the optimal perturbation magnitude and the direction of the membership signal vary across model–dataset pairs. For many settings, a previously unstudied low-noise regime produces stronger separation than large perturbations, with members exhibiting lower divergence than non-members and thereby reversing the conventionally observed score direction. Other settings retain the conventional direction, demonstrating that neither a fixed perturbation magnitude nor a predetermined score orientation is universally reliable. Motivated by this finding, we propose COMPASS, an adaptive gray-box MIA that calibrates both the perturbation magnitude and membership-score direction using a loosely distribution-matched, known non-member reference set. We instantiate COMPASS using Rényi-normalized KL divergence and Rényi divergence, with max-k aggregation over the most informative image-logit positions. Across LLaVA v1.5–7B, MiniGPT-4, and HuluMed-32B, spanning pretraining, instruction-tuning, and medical-image datasets, COMPASS consistently outperforms baselines. Under general-domain evaluations, COMPASS achieves AUCs of up to 0.88. On the distribution-matched medical evaluation, it reaches 0.64 AUC, while the strongest baseline remains at chance level. These results establish full-range adaptive perturbation analysis as an effective approach to VLLM membership inference.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.