VTEGuard: Proactive Defense against Diffusion-Based Visual Text Editing
Abstract
Latent diffusion models enable realistic visual text editing but also expose publicly shared images to unauthorized textual manipulation. A central challenge in proactive protection is to prevent the intended text edits while preserving the visual quality of the released image. However, overall visual degradation of the edited output does not necessarily prevent the target text from being correctly rendered, motivating a focused investigation of protection against diffusion-based visual text editing. We propose VTEGuard, a proactive defense that couples full-image frequency-domain perturbation optimization with variational autoencoder (VAE) posterior-variance disruption. Specifically, VTEGuard parameterizes protective perturbations in the full-image YCbCr discrete cosine transform (DCT) domain, enabling globally structured modifications under a bounded perturbation constraint. We further introduce an Entropic-Smoothed Linear (ESL) objective that combines global posterior-variance expansion with adaptive emphasis on low-variance latent dimensions. This objective maintains dense, non-saturating gradients in variance space and guides the frequency coefficients through the VAE encoder and inverse frequency transform. Experiments on VTEGuardBench and SDEdit, DiffEdit, DreamBooth-based personalization, and diffusion-based inpainting show that VTEGuard achieves the best performance among the evaluated methods across multiple metrics. These results demonstrate a favorable trade-off between protection effectiveness and protected-image fidelity, while highlighting its broad applicability beyond visual text editing.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.