Beyond Visual Artifacts: Knowledge-Intensive Detection of AI-Generated Images
Abstract
Detectors that rely on pixel-level artifacts or semantic visual flaws can struggle with increasingly realistic AI-generated images (AIGIs). Prior AIGI detectors have not systematically exploited a complementary signal: knowledge inconsistencies between implicit visual claims and real-world facts. Such inconsistencies can arise from generator hallucinations or deliberate creator choices. To evaluate this signal, we introduce K-AIGI-Bench, to our knowledge the first benchmark specifically designed for knowledge-intensive AIGI detection. It contains carefully human-curated, high-quality images produced by a state-of-the-art image generator, each with a deliberately embedded knowledge inconsistency, alongside topic-matched real images collected from the Internet. To address this task, we further propose K-VERA (Knowledge-grounded Visual Evidence Retrieval and Adjudication), a training-free, retrieval-augmented multi-agent framework. It identifies the scene and entities, turns implicit visual claims into independent verification tasks, and retrieves webpages and reference images to verify them. It then checks source support and resolves conflicting findings to produce an evidence-grounded verdict. Experiments on K-AIGI-Bench show that K-VERA outperforms all evaluated dedicated AIGI detectors by over 15% in balanced accuracy, while producing higher-quality, evidence-grounded explanations. Combining K-VERA with a pixel-level detector further improves performance on three in-the-wild benchmarks. These results establish knowledge-intensive detection as an important enhancement for detecting increasingly realistic AIGIs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.