Frequency Attribution Mask: Toward Detector-Adaptive Robustness in AI-Generated Image Detection
Abstract
AI-generated image (AIGI) detectors report high accuracy on benchmark splits, yet collapse on unseen generators and after routine post-processing such as JPEG compression or resizing, exposing a reliance on shortcut cues that do not transfer. Existing remedies name a specific source-level bias and patch it with a deterministic operation at the data, architecture, or weight level, leaving every other bias untouched and tying the fix to a single detector setup. To address this, we propose *Frequency Attribution Mask* (FAM), an attribution-guided mask-tuning framework that intervenes at each detector's own input sensitivity, the point at which any such bias actually manifests. FAM derives a per-sample frequency-domain attribution from the detector's input gradient, masks the highest-attribution components of the input spectrum with class-asymmetric thresholds, and finetunes on the masked reconstructions. We apply FAM to two pretrained ResNet-50 detectors, one of which already carries a pixel-level alignment fix, and show that it consistently improves post-processing robustness over both base detectors and recent shortcut-mitigation baselines, with the largest gains where residual shortcut reliance is most severe.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.