acceptodds
Under review as a conference paper at ICLR 2027

Not All Channels Are Equal: Perturbation-Invariant Channel Selection for Robust AI-Generated Image Detection

Abstract

AI-generated image detectors often lose accuracy after compression, blurring, noise, or resizing. We ask whether this sensitivity is concentrated in particular feature channels of a pretrained detector, and on a forensic-trained CNN we find that it is: the channels whose spatial activation patterns survive perturbation are not the ones the classifier relies on most. We propose PICS (Perturbation-Invariant Channel Selection), which scores each channel by the Pearson correlation between its clean and perturbed activation maps, keeps the union of the highest-ranked channels across perturbations, and retrains the linear head on the selected channels. Under a common training protocol, PICS improves robustness not only beyond retraining the head with augmentation alone, but also beyond magnitude-based, random, attention-based, and weight-clamping alternatives, as well as end-to-end fine-tuning of the backbone. The recipe applies to CNN and ViT detectors, with the largest gains over the released detectors on forensic-trained backbones whose robust and discriminative channels are most separated, and we evaluate it on cross-dataset benchmarks and social-media platforms. These results support perturbation stability as a practical criterion for selecting the features of forensic detectors.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.