Rethinking Image Immunization for Transferable and Purification-Robust Protection
Abstract
Image immunization has been actively studied to protect images from unauthorized diffusion-based manipulation. For a practical deployment, however, image protection requires two key properties: black-box transferability to unknown generative models and robustness against perturbation purification, which have largely been studied along separate directions. In this paper, we present a unified immunization framework that jointly targets both requirements. Our controlled analysis identifies early VAE encoder representations as a promising common attack surface: they exhibit strong cross-model correspondence, and perturbations optimized at these stages better retain their induced latent shifts after purification. Based on these observations, we directly disrupt early VAE encoder representations, rather than relying on architecture-specific denoisers or explicit purification-aware designs. We further introduce a relation-disruption objective that disrupts the internal cross-channel structure of early representations. Extensive experiments demonstrate strong black-box transferability across heterogeneous diffusion-based generative models within multiple downstream manipulation tasks, and robustness against diverse purification techniques, with low computational overhead.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.