acceptodds
Under review as a conference paper at ICLR 2027

The Search for Statistical Indistinguishability Between Real and Fake Images

Abstract

Images produced by generative models appear strikingly similar to real ones. In this work, we study whether perfection has been reached by generated content in terms of matching the properties of real data. We train deep neural networks to separate real data from fake and observe their test accuracies. Only a chance performance has any scope of proving statistical indistinguishability. However, since generative models are known to sometimes introduce small errors, e.g., low level checkerboard artifacts, we create progressive abstractions of the original real/fake images to potentially bring them to a space where those artifacts get wiped away. These include operations such as blurring, downsizing, cropping, converting RGB images into edge/semantic maps, etc. To our surprise, almost none of the abstractions are able to make fake data indistinguishable from real. The only cases when neural network based classifiers fail to separate the two categories is when the data is abstracted to such an extent that the original content seems mostly lost to the human eye, e.g., downsizing a RGB image of human face to resolution. We observe this trend across four different datasets of varying complexity – MNIST, CelebA, FFHQ and GenImage – and across various types of generative models – GANs, VAEs, diffusion models. This shows that the nature of fake artifacts extends all the way from low level to very high level in images. The simple implication is that fake images are much more different from real than we are used to think.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.