acceptodds
Under review as a conference paper at ICLR 2027

Residual Family Augmentation for Detecting Images from High-Fidelity Autoencoders

Abstract

AI-generated image (AIGI) detection has advanced rapidly, yet on the newest text-to-image models even robust detectors degrade sharply. In this paper, we diagnose the failure with a Residual Scaling Sweep, which scores a detector as the reconstruction residual is scaled, and identify an overlooked patch-size mismatch. The sweep shows that detectors trained with existing recipes respond only near the magnitude and the sign of their training residual, and they fail to keep up with recent generators. We therefore propose Residual Family Augmentation: the residual is added back to the photograph within a bounded range of magnitudes and at both signs, all labeled fake, plus a second reconstruction pass. The detector then no longer depends on the magnitude or the sign of one residual and responds across that range, which includes the weaker residuals of autoencoders never seen in training. We further observe that the autoencoders of recent generators leave a reconstruction residual, the difference between a photograph and its reconstruction, that is weaker and concentrated at finer spatial scales than the residual of earlier autoencoders, scales that the pretrained filters of vision backbones attenuate. Notably, we find that simply halving the patch of the pretrained backbone, which adds no parameter, restores the residual to scales the backbone retains, and the augmentation reaches its full gain only under the resampled patch. Trained on a single, older autoencoder, our detector leads seven public detectors by five points or more in mean balanced accuracy over eleven sets from the newest generators, while remaining competitive with the best baseline across public benchmarks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.