EFG: Encoder-Free Forensic Guard for Synthetic Image Detection and Explanation
Abstract
AI-generated image detection has recently benefited from vision-language models (VLMs), which provide natural-language explanations to justify their judgments. However, existing VLM-based detectors heavily rely on visual encoders, which compress images into high-level semantic features, inevitably discarding low-level synthetic artifacts critical for forensic discrimination. Consequently, this compression creates a fundamental bottleneck, with detection-relevant signals largely lost before the model even begins reasoning. In this work, we propose EFG (Encoder-Free Forensic Guard), a specialized VLM that captures comprehensive forensic signals without visual encoders. Specifically, EFG converts the input image into token sequences via a lightweight patch embedding layer and directly feeds these raw pixel-level tokens into the large language model. This design preserves complete visual signals without encoder-induced information loss, while adaptively fusing low-level forensic traces with high-level semantics across layers. Beyond detection, EFG can produce natural-language explanations to justify its forensic judgments. Extensive experiments show that EFG achieves state-of-the-art performance across three AI-generated image detection benchmarks. Our work highlights the potential of encoder-free, pixel-space learning for VLM-based forensic systems. Code will be open-sourced upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.