When the Service Is the Adversary: Cue-Driven Editing Against VLM Attribute Inference
Abstract
When users submit images to closed-source vision-language models (VLMs) for everyday tasks, the model processes not only what the user explicitly asks, but also what the image passively reveals—age, gender, ethnicity, occupation, and other sensitive attributes the user never intended to share. We formalize this threat as Service-Time Private Attribute Inference (ST-PAI) and show that existing perturbation-based defenses are structurally insufficient under it: the target model is closed-source and controls its own preprocessing pipeline, stripping pixel-level perturbations before inference is ever performed. We instead build defense on top of semantic image editing, with attribute-related visual cues sequentially edited out of the image according their importance, producing a discrete privacy-utility frontier that trades utility for protection in a controlled manner. Experiments on two geolocation benchmarks show that semantic editing remains robust to upstream purification, whereas perturbation-based protection collapses, supporting our argument that inferential evidence must be removed at the semantic level. We further repurpose existing VQA tasks to analyze how the utility cost of protection depends on the downstream task, and construct SynthPAI-Visual, a synthetic multi-attribute benchmark with automatically generated attribute annotations that requires no real private data.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.