Beneath Semantics: Signal‑axis Residual Fingerprint for AI‑Generated Image Detection
Abstract
AI-generated image detectors perform well on aggregate benchmarks but exhibit marked declines in performance under distribution shift—whether from new collection pipelines for familiar generators or from entirely unseen generators. We examine image authenticity along two complementary evidence axes: semantic representations, which provide detection cues through pretrained image representations, and signal statistics, which capture local dependencies and spectral organization in residuals. To address both collection shift and unseen-generator shift, we introduce Signal-axis Residual Fingerprint (SRF), a compact detector built around a fixed representation of Spatial–Spectral Microstructure. Rather than learning dataset- or generator-specific patterns, SRF uses six fixed resampling operators and an interpretable statistic bank to characterize local dependencies and spectral organization in residuals—traces intrinsic to the generation pipeline rather than specific to a particular collection protocol or generator. The complete five-block configuration, SRF-E, contains only 0.28M fitted parameters, requires no pretrained backbone, and achieves a mean standalone AUROC of 0.9150 across 48 benchmark–generator groups from four benchmark suites. Crucially, its scalar score can be fused with the 12 evaluated frozen-backbone probes to recover performance under both distribution shifts. For the strongest CLIP probe, mean AUROC increases by approximately 2.6 percentage points, worst-group AUROC rises from 0.691 to 0.817, and mean AUROC on the unseen-generator tier improves from 0.855 to 0.899. Controlled transformation experiments show that periodic residual correlations and spectral energy statistics provide processing-related evidence that the current CLIP detection scores do not fully exploit.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.