acceptodds
Under review as a conference paper at ICLR 2027

Prompt Prototype Learning (PPL): Rethinking Text Supervision for AI-Generated Image Detection

Abstract

Contrastive vision-language models (VLMs) have shown promise for generalizable AI-generated image (AIGI) detection, yet existing approaches rely on fixed single-token class labels or per-sample captions, a brittle form of supervision that fails to capture diverse cues to generalize across unseen generators. We propose Prompt Prototype Learning (PPL), a novel framework that replaces fixed class tokens with ensemble-averaged prototype embeddings. For each class, we construct multiple prompts capturing complementary aspects of visual realism, including texture, boundary consistency, and high-level semantics. These prompts are encoded and aggregated into a normalized class prototype, providing richer and more stable contrastive supervision. Training images are aligned with these prototypes, improving generalization to variations in generation artifacts. We apply lightweight LoRA to the vision encoder and train exclusively on ProGAN with four object categories. Despite this limited training distribution, PPL generalizes strongly to unseen generators in the UniversalFakeDetect (97.3% mACC and 99.8% mAP) benchmark, as well as across four additional datasets covering recent generators and real-world transmitted images (94.1% mACC and 97.7% mAP), spanning GANs, diffusion models, and commercial tools, surpassing all prior methods. Ablation studies confirm that PPL is the primary driver of performance gains, with texture, boundary, and geometric prompts each contributing complementary signals.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.