Seeing Is Not Recognizing: Exact Target Identification as the Bottleneck for Toxicity Detection in AI-Generated Visual Illusions
Abstract
Controllable diffusion models can embed specified targets into natural scenes with benign surface content to create visual illusions. Toxic visual illusions pose new challenges for multimodal content moderation. Existing benchmarks have limited coverage and lack non-toxic controls and non-illusion images, while end-to-end evaluation makes it difficult to localize errors. We introduce ToxicIllusion, a human-annotated benchmark comprising 14,940 images generated from 166 toxic and non-toxic targets across four target types: English text, Chinese text, gesture, and symbol. We propose a three-stage evaluation framework covering illusion perception, target identification, and toxicity classification to localize errors across these stages. We examine the relationship between target identification and toxicity detection through output state decomposition and conditional toxicity analysis. We further analyze visual representations and relevance maps to examine the representational basis of target identification failures and models' spatial responses to surface scenes and targets. Our results show that current vision-language models do not reliably connect illusion perception, target identification, and toxicity classification. Target identification is a key bottleneck in toxicity detection, and correct target identification is consistently associated with more reliable toxicity classification. The visual representations of pretrained visual encoders are dominated by surface scenes. Relevance maps further show that the evaluated open-weight models primarily attend to surface scenes rather than targets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.