Learning Transferable Semantic Image Edits for VLM Resource-Exhaustion Attacks
Abstract
The inference cost of a vision-language model (VLM) service can be amplified by an adversary-supplied image that induces unusually long outputs under a fixed benign prompt. blackHowever, the effectiveness of pixel-space attacks does not consistently generalize across VLM architectures. To improve transfer, we introduce Visual Editing for Token Amplification (VETA), an image-only resource-exhaustion attack that uses reinforcement learning (RL) to train an image-editing planner from the returned responses of one surrogate VLM. The planner specifies changes to visible objects, text, and scene relations; a frozen diffusion editor executes these instructions while preservation rewards encourage retention of the original scene. Two-stage RL first learns reliable output-lengthening edits and then rewards exceptionally long and limit-reaching outputs. All compared methods generate four candidates per source image and use the same evaluation rules. Across six target VLMs, VETA achieves macro amplification when averaging the four candidates and when selecting using only surrogate feedback, without target-side selection. Separately, the maximum over four target-model responses reaches , measuring the longest observed output rather than average request cost. Same-editor controls and ablations assess the contribution of learning, while controlled serving measurements of target-selected candidates show increased latency and energy use and reduced throughput. Learned semantic image editing thus enables transferable resource amplification.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.