The Evidence Within: Tracing and Recovering Generative Evidence for Generalizable AI-Generated Image Detection
Abstract
AI-generated image detection has become increasingly important as synthetic images rapidly enter real information ecosystems. Recent visual foundation models (VFMs) equipped with lightweight linear heads have emerged as strong baselines for this task. However, like many previous detectors, they still suffer from fake-side generalization failures under unseen or shifted generators, including cross-mechanism transfer, reconstruction-style generation, and localized editing. We investigate these failures from a representation perspective: when the final image-level prediction fails, has generative evidence disappeared from the model, or does it remain encoded elsewhere in its internal representations? To study this question, we propose Mechanism-Aligned Evidence Probing (MAEP), a role-adaptive diagnostic framework that turns mechanism-specific linear heads into evidence probes and traces generation-related responses across representation depth and token types in frozen DINOv3. MAEP reveals that generative evidence is not uniformly expressed in a single representation. Its strength and persistence vary across generation mechanisms, layers, and token types: images predicted as real by the final image-level decision can still retain measurable fake-side evidence in intermediate representations or patch tokens, and the most discriminative evidence often emerges before the final layer. Motivated by this finding, we propose MIRAGE, Multi-dimensional Internal Rarity Aggregation for Generalizable Evidence, which reuses a trained linear head as an internal evidence probe to recover generation traces overlooked by the image-level decision. MIRAGE measures fake-side evidence across selected internal dimensions, converts it into empirical rarity with respect to clean-real calibration data, and aggregates complementary rarity signals through calibrated empirical min-p. Without additional backbone training, MIRAGE consistently improves generalization across diverse benchmarks and challenging distribution shifts, outperforming all compared methods on major benchmarks. In particular, it improves average balanced accuracy from 81.8% to 97.3% on AIGCDetectionBenchmark and from 91.4% to 94.3% on Chameleon, while yielding substantial gains on localized and reconstruction-style fakes. These results suggest that generative evidence follows structured, mechanism-dependent trajectories within VFM representations, and that recovering such internal evidence provides an effective route toward more generalizable AI-generated image detection. The code will be made publicly available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.