EviSlot: Evidence Slot-Guided Vision-Language Learning for Generalizable AI-Generated Image Detection
Abstract
AI-generated images have become increasingly realistic and visually indistinguishable from real images, making generalizable AI-generated image detection a challenging task. Existing vision-language detectors usually rely on global visual representations and deterministic textual prompts, which may overlook localized high-frequency artifacts and provide limited flexibility under cross-generator distribution shifts. To address this issue, we propose EviSlot, an evidence slot-guided vision-language learning framework that jointly models global semantics, local forensic evidence, and diverse textual hypotheses. Specifically, we employ DINOv3 as the visual encoder to extract both CLS and patch-level features, and introduce a DCT-guided evidence mining module to identify frequency-abnormal regions with potential artifact cues. A learnable Evidence Slot module then aggregates local artifact cues from patch tokens with frequency-aware attention guidance and fuses them with the global representation for robust visual evidence modeling. On the textual side, frequency-aware probabilistic prompting generates multiple real/fake prototypes from shared, image-specific, and class-specific prompt components. The final prediction is obtained by image-text matching and prompt-ensemble inference. The proposed framework provides a compact, flexible, and robust detection paradigm for generalizable AI-generated image detection.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.