ScaleTrojan: Uncovering Backdoor Vulnerability in Visual Autoregressive Text-to-Image Generation
Abstract
Visual autoregressive (VAR) models have recently emerged as a competitive al- ternative to diffusion models for text-to-image generation, yet their security prop- erties remain largely unexplored. In this work, we systematically investigate backdoor vulnerability in VAR-based text-to-image generation. We first develop ScaleTrojan, a backdoor attack framework that combines multi-scale target learn- ing with clean prediction regularization against the frozen pretrained model. Our ablations reveal two key findings: full-scale supervision provides substantially stronger target control than partial-scale supervision, while removing clean pre- diction regularization retains high attack success but severely degrades clean gen- eration. We then examine backdoor injection subspaces and find that broader parameter adaptation does not consistently improve the attack–utility trade-off, with cross-attention providing an effective injection interface. As an initial explo- ration of backdoor detection, we propose a method based on the consistency of predictions across visual scales under textual perturbations. Experiments on dif- ferent VAR backbones demonstrate that ScaleTrojan achieves nearly 100% attack success rate (ASR) while largely preserving clean generation quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.