acceptodds
Under review as a conference paper at ICLR 2027

The Evidence Within: Tracing and Recovering Generative Evidence for Generalizable AI-Generated Image Detection

Abstract

AI-generated image detection has become increasingly important as synthetic images rapidly enter real information ecosystems. Recent visual foundation models (VFMs) equipped with lightweight linear heads have emerged as strong baselines for this task. However, like many previous detectors, they still suffer from fake-side generalization failures under unseen or shifted generators, including cross-mechanism transfer, reconstruction-style generation, and localized editing. We investigate these failures from a representation perspective: when the final image-level prediction fails, has generative evidence disappeared from the model, or does it remain encoded elsewhere in its internal representations? To study this question, we propose Mechanism-Aligned Evidence Probing (MAEP), a role-adaptive diagnostic framework that turns mechanism-specific linear heads into evidence probes and traces generation-related responses across representation depth and token types in frozen DINOv3. MAEP reveals that generative evidence is not uniformly expressed in a single representation. Its strength and persistence vary across generation mechanisms, layers, and token types: images predicted as real by the final image-level decision can still retain measurable fake-side evidence in intermediate representations or patch tokens, and the most discriminative evidence often emerges before the final layer. Motivated by this finding, we propose MIRAGE, Multi-dimensional Internal Rarity Aggregation for Generalizable Evidence, which reuses a trained linear head as an internal evidence probe to recover generation traces overlooked by the image-level decision. MIRAGE measures fake-side evidence across selected internal dimensions, converts it into empirical rarity with respect to clean-real calibration data, and aggregates complementary rarity signals through calibrated empirical min-p. Without additional backbone training, MIRAGE consistently improves generalization across diverse benchmarks and challenging distribution shifts, outperforming all compared methods on major benchmarks. In particular, it improves average balanced accuracy from 81.8% to 97.3% on AIGCDetectionBenchmark and from 91.4% to 94.3% on Chameleon, while yielding substantial gains on localized and reconstruction-style fakes. These results suggest that generative evidence follows structured, mechanism-dependent trajectories within VFM representations, and that recovering such internal evidence provides an effective route toward more generalizable AI-generated image detection. The code will be made publicly available.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.