Learning Adaptive Evidence Selection for Generalizable AI-Generated Image Detection
Abstract
Generalizable AI-generated image (AIGI) detection aims to identify synthetic images and generalize to unseen generators.To improve the generalization ability, recent detectors increasingly adopt Vision Foundation Models (VFMs).However, most VFM-based detectors use the VFM as the image representation, leaving two issues unresolved.First, the global token is dominated by high-level semantic content, which can induce reliance on semantic content rather than generation-related forensic cues.Second, the global token focuses on a few patches, leaving complementary evidence in other patches underused.To address these issues, we propose the Adaptive Evidence Selection (AES) framework for generalizable AIGI detection.To reduce the influence of semantic content, we propose Cluster-centered Bilinear Aggregation (CBA), which suppresses shared semantic content and models local features within each cluster.To exploit forensic cues distributed across regions, we propose Determinantal Evidence Selection (DES), which selects cluster features by jointly evaluating task relevance and feature diversity. Experiments on four benchmarks demonstrate the effectiveness of our proposed method.Our source code will be released soon.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.