acceptodds
Under review as a conference paper at ICLR 2027

ScaleTrojan: Uncovering Backdoor Vulnerability in Visual Autoregressive Text-to-Image Generation

Abstract

Visual autoregressive (VAR) models have recently emerged as a competitive al- ternative to diffusion models for text-to-image generation, yet their security prop- erties remain largely unexplored. In this work, we systematically investigate backdoor vulnerability in VAR-based text-to-image generation. We first develop ScaleTrojan, a backdoor attack framework that combines multi-scale target learn- ing with clean prediction regularization against the frozen pretrained model. Our ablations reveal two key findings: full-scale supervision provides substantially stronger target control than partial-scale supervision, while removing clean pre- diction regularization retains high attack success but severely degrades clean gen- eration. We then examine backdoor injection subspaces and find that broader parameter adaptation does not consistently improve the attack–utility trade-off, with cross-attention providing an effective injection interface. As an initial explo- ration of backdoor detection, we propose a method based on the consistency of predictions across visual scales under textual perturbations. Experiments on dif- ferent VAR backbones demonstrate that ScaleTrojan achieves nearly 100% attack success rate (ASR) while largely preserving clean generation quality.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.