Probing AI-Generated Image Detectors via Black-Box Anti-Forensics on Semantic and Artifact Cues
Abstract
Recent advances in generative models have made AI-generated images increasingly difficult to distinguish from real ones, motivating detectors that exploit semantic representations from vision foundation models (VFMs) and synthesis-specific artifacts left by generative models. However, it remains unclear whether these cues capture intrinsic evidence of forgery or merely reflect incidental characteristics of the generation process. In this work, we investigate this question through systematic anti-forensic probing, finding that detectors can be substantially compromised by independently manipulating these two types of evidence: the semantic representations used for discrimination can be steered toward authentic semantics, while synthesis-specific artifacts can be suppressed without noticeably altering image content. Based on these findings, we propose CRAS, a black-box anti-forensic framework that jointly targets semantic and artifact evidence. Specifically, a cross-modal reconstruction module steers forged images embeddings toward authentic text anchors in the VFM feature space, followed by an artifact suppression module that reduces the synthesis-specific artifacts left by the reconstruction process. Extensive experiments show that CRAS achieves the best anti-forensic performance among existing methods while preserving high visual fidelity. Our results suggest that current AIGI detectors rely heavily on both types of evidence that are not invariant to carefully controlled transformations, highlighting the importance of learning manipulation-invariant forensic evidence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.