Language Priors without Language: A Frozen LLM Layer for Generated Image Detection
Abstract
Generated image detection increasingly relies on frozen vision-language models such as CLIP, whose features generalize well across generators but are learned through global image-text alignment and remain limited in modeling semantic relationships among image regions and structural consistency. As modern generators suppress low-level artifacts, high-level cues such as semantic and structural anomalies become decisive. We observe that large language models (LLMs) acquire, through autoregressive pretraining, precisely the ability that CLIP lacks: organizing semantic relationships and contextual dependencies among tokens. We therefore propose to treat a single frozen decoder layer of a pretrained LLM as a semantic relation enhancement module for visual tokens. The layer re-encodes the CLIP tokens within the relational structure learned from language, without any language prompts or textual inputs, while both CLIP and the LLM remain frozen and only about 4.3M parameters are trained. Trained on a single generator, our detector reaches 91.55% accuracy and 99.38% average precision across 13 GAN and diffusion generators, surpassing the strongest prior detector by 3.05 accuracy points, and it generalizes to recently released diffusion and autoregressive generators with 98.43% accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.