FakeCueS: Learning Scene-Conditioned Cues for AI-Generated Image Detection
Abstract
Advances in generative models enable high-fidelity image synthesis across diverse scenes, characterized by visual style and subject content, challenging reliable AI-generated image detection. Recent multimodal detectors identify visual evidence for their decisions but do not explicitly model its scene-dependent diagnostic value: textures or structures that suggest AI generation in photographs may reflect ordinary artistic expression in paintings, potentially leading to misclassification. To address this limitation, we introduce FakeCueS, a framework that uses scene context to guide forensic cue generation and evaluation. Predicted scene representations guide the model to generate context-specific cues. An online cue evaluator verifies visual support and groups semantically related cues to assess their source association within each scene. These assessments provide set-level rewards for group-relative policy learning alongside image-level supervision, without reference rationales or region annotations. We further introduce RedScenes, a dataset of 168,688 social-media images spanning 70 scenes. FakeCueS achieves 91.53% accuracy on RedScenes, exceeding the strongest evaluated zero-shot baseline by 12.25%. Without target-domain fine-tuning, it achieves 94.49% and 83.09% accuracy on AnomReason-Deepfake and Chameleon, respectively.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.