acceptodds
Under review as a conference paper at ICLR 2027

ADVO: Protecting Images from Exploitation by Closed-Source Multimodal Systems

Abstract

Closed-source multimodal systems can describe, edit, and imitate images through natural-language instructions, making unauthorized exploitation of personal images increasingly accessible. Protecting images against these systems is challenging because their internal architectures and parameters are unavailable, while existing image-protection methods are typically optimized for known models and transfer poorly to closed-source targets. In this work, we propose ADVO, a transferable image protection framework that protects images against closed-source systems by adding a small, visually subtle perturbation with the assistance of open-source proxy models. ADVO combines two complementary protection mechanisms: refusal induction, which encourages the target system to reject the protected image, and source-semantic disruption, which suppresses exploitable semantics when refusal is not triggered. To make these mechanisms transferable to inaccessible closed-source systems, ADVO adopts a two-stage framework. In the first stage, ADVO extracts transferable refusal features from policy-violating reference images (e.g., images containing NSFW content) while filtering out irrelevant visual information. In the second stage, ADVO assigns the two mechanisms to different image regions to reduce their interference under a limited perturbation budget. ADVO adaptively balances the contributions of different proxy models during optimization to improve transferability across heterogeneous closed-source systems. We evaluate ADVO across multiple image-exploitation tasks, including image description, editing, and imitation. On image description across six closed-source systems, including GPT-5.6 and Gemini 3.8, ADVO achieves the highest average protection success rate (PSR) of among all evaluated methods, with a maximum refusal rate of . Even on the more permissive Grok 4.6, ADVO achieves PSR despite only refusals, illustrating the complementarity of refusal induction and source-semantic disruption. Results on image editing and imitation further suggest that the same protection extends to downstream image-exploitation tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.