Provenance is not Universal: Discovering and Distilling Forensic Subspaces in Frozen Vision-Language Models
Abstract
AI-generated image detectors often achieve strong in-distribution performance but degrade on unseen generators and datasets, raising the question of whether pretrained vision-language models (VLMs) encode more transferable provenance signals. Although frozen features from foundation models have shown promising performance for synthetic image detection, it remains unclear how such provenance information is organized within VLM representations. We address this gap by studying prompt-conditioned hidden states from seven frozen VLMs. We introduce Behavioral Provenance Signatures (BPS) to characterize separability, generator-conditioned forensic subspaces, cross-domain transferability, and distillability. We observe that frozen representations consistently support strong in-distribution separation but exhibit substantial variation across models and source generators. Generator-conditioned probes have partially aligned yet distinct discriminant directions, showing that strong separability does not imply a universal provenance axis. Distillation is similarly heterogeneous across teachers, target domains, and student architectures. Nevertheless, the resulting image-only models eliminate VLM execution and prompt processing at test time, achieving more than a measured inference speed-up. These results distinguish the availability of provenance information from its robustness and deployability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.