acceptodds
Under review as a conference paper at ICLR 2027

Image Membership Inference in Fine-Tuned Medical Vision Language Models via Decoding Surfaces

Abstract

Membership inference attacks (MIAs) are widely used to audit training-data leakage in fine-tuned medical vision–language models (VLMs). Existing VLM MIAs typically commit to a fixed decoding configuration or score family before evaluation, even though stochastic generation behavior can vary substantially with the inference policy. This restricted view may therefore miss stronger membership evidence elsewhere in the observable response space. We propose the Decoding-Surface Attack (DSA),, a calibrated black-box MIA that models the multi-temperature response surface at each output budget by combining within-policy consistency, cross-temperature geometry, and response behavior. Beyond attack measurement, DSA uses the resulting leakage score to guide surface-access restriction and privacy-aware checkpoint selection, forming an audit-to-mitigation workflow without claiming formal privacy guarantees. Extensive evaluations across medical VQA benchmarks and model families compare DSA with fixed-policy, partial-surface, and recent VLM MIA baselines. DSA achieves 0.811 macro AUC in our primary evaluation, improving over fixed-policy MIA by 25.5 AUC points and the strongest reproduced VLM MIA baseline by 18.1 points. Comprehensive component ablations and same-pool controls further verify the contribution of complementary surface evidence and rule out the tested dataset shortcuts.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.