acceptodds
Under review as a conference paper at ICLR 2027

One Man's Cure Is Another's Poison: Turning JPEG Compression Into Backdoor Triggers for Vision-Language (Action) Models

Abstract

JPEG compression is ubiquitous in practical vision-language model (VLM) pipelines for storage, transmission, and input preprocessing, and has also been evaluated as a lightweight defense against adversarial perturbations. Existing studies largely assume that compression suppresses malicious behavior or represents a distortion that adversarial examples should survive. In this work, we find that such a defense may instead be exploited by the adversary as a trigger for attacks, i.e., the input appears benign but induces attacker-specified behavior after JPEG compression in vision-language(-action) tasks. We propose JPEG-triggered Adversarial attaCK (JACK), a two-stage method that achieves pre-compression stealthiness and post-compression attack effectiveness without modifying the victim model or the JPEG codec. JACK first constructs a candidate image whose JPEG-compressed version induces the target behavior, and then restores benign behavior for the uncompressed counterpart within the same JPEG quantization cell, thereby preserving the attack-inducing compressed representation. Extensive experiments on targeted image classification, visual question answering, multimodal jailbreaking, and multi-view Vision-Language-Action (VLA) policy hijacking validate the effectiveness of the proposed method.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.