Cross-Layer Relational Representation Modeling for Few-Shot AI-Generated Image Detection
Abstract
AI-generated images have become increasingly realistic, raising urgent concerns about content authenticity and misuse, making reliable detection under unseen generators increasingly important. Since collecting large-scale training data for every emerging generator is impractical, this challenge motivates few-shot AI-generated image (AIGI) detection, where a detector must adapt to novel generators from only a handful of labeled examples. In this work, we identify generator-dependent cross-layer relational patterns in pretrained CLIP representations as a generalizable cue for AIGI detection. Specifically, real images and images from different generators exhibit distinct cross-layer behaviors in token representations and class-patch relational patterns. Motivated by this analysis, we propose a Cross-Layer Relational Module (CLRM), which constructs relational representations through bidirectional cross-attention between class tokens and patch tokens from intermediate and final layers. CLRM then transforms such semantic-spatial correspondence patterns across layers into metric-space embeddings and combines them with prototypical learning for few-shot adaptation. Extensive experiments demonstrate that CLRM achieves state-of-the-art performance in few-shot classification and generator source attribution, while providing strong generalization and robustness.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.