CLARE: Continuous Latent Reasoning for AIGC Image Detection
Abstract
As AI-Generated Content (AIGC) advances rapidly, the growing risk of its malicious exploitation creates an urgent need for reliable AIGC image detection. While recent approaches based on Multimodal Large Language Models (MLLMs) have achieved promising detection performance, their reliance on explicit textual chain-of-thought (CoT) introduces two fundamental limitations. First, autoregressively decoding continuous representations into discrete text tokens creates an information bottleneck, compromising rich visual cues. Second, as generated images become increasingly realistic, constructing high-quality and hallucination-free textual reasoning chains becomes substantially more difficult. To overcome these limitations, we propose CLARE, a Continuous Latent Reasoning framework for AIGC image detection. CLARE conducts forensic reasoning entirely within the continuous latent space of MLLMs by directly feeding the last hidden state of each step as the input representation for the next step, thereby bypassing discrete text token generation. Furthermore, we pretrain a task-specific vision encoder on AIGC detection and inject its representations into the latent reasoning process, effectively equipping the general multimodal backbone with forensic sensitivity. Extensive experiments across multiple challenging benchmarks demonstrate that CLARE achieves superior detection performance and exhibits strong cross-generator generalization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.