PRISM: Post-hoc Recognition of Image Source Models
Abstract
Image generative models (GMs) now produce photorealistic content that is virtually indistinguishable from natural images. This creates a critical need for reliable provenance to distinguish between synthetic vs authentic images and attribute generated content to its source. While diffusion models (DMs) have been the dominant paradigm for visual content generation, recent image autoregressive models (IARs) match DMs' generation quality and substantially improve the sampling speed by adapting the next-token prediction paradigm from large language models. We propose a unified post-hoc attribution method that identifies the source GM of a given image by exploiting model-specific artifacts (fingerprints) that arise in both DMs and IARs. Our key insight is to combine the fingerprints from the two main constituent parts of GMs: (1) the autoencoder that maps between pixel and latent spaces and (2) the latent generator that defines the image distribution in DMs and produces visual tokens in IARs. In the first stage, we identify the autoencoder family using reconstruction and quantization errors. In the second stage, we attribute the image to the latent generator within the identified model family using conditional log-likelihoods for DMs and token log-probabilities for IARs. Across diverse model families and versions, our two-stage framework enables accurate model attribution, providing a unified approach to the provenance of synthetic images.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.