Learning Generalizable Model Fingerprints for Attribution of Unseen Generative Models
Abstract
In this work, we propose a synthetic image attribution method based on generative model fingerprints that requires no synthetic images from real-world generative models during training. Instead, we train the fingerprint extractor entirely on a large and controllable bank of reconstruction models spanning diverse architectures and independently perturbed parameterizations, using the fingerprints embedded in their reconstructed images as supervisory signals. The proposed framework learns content-invariant fingerprints that distinguish not only different architectures but also different parameterizations of the same architecture, enabling direct attribution to previously unseen generative models. We further adopt an agent-guided collaborative training strategy to adaptively coordinate sampling, local evidence selection, loss configuration, frequency scheduling, and model-bank expansion during training. We evaluate the learned fingerprints on 26 unseen source models collected from four public benchmarks. Our method achieves a macro-average verification AUC of 97.01%, outperforming the strongest baseline by 5.46 percentage points, and shows clear separation between intra-source and inter-source fingerprints. These results demonstrate that fingerprints learned solely from controlled reconstruction models can generalize effectively to diverse unseen generative models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.