Unlocking Native Forensic Representation in Multimodal Large Language Models
Abstract
Multimodal Large Language Models (MLLMs) have emerged as a promising paradigm for AI-generated image (AIGI) detection by leveraging high-level semantic understanding. However, reliable detection also depends on subtle low-level forensic evidence, motivating growing efforts to equip MLLMs with stronger forensic perception. This work demonstrates that strong forensic discrimination can be achieved using only the native vision representations of pretrained MLLMs. We introduce **Native Forensic Connection (NFC)**, a remarkably simple framework that directly connects native forensic representations to the language model. NFC is inspired by our observation that low-level forensic information is already highly discriminative in intermediate vision representations, but becomes substantially less accessible along the standard vision-language pathway. With NFC, MLLMs can access native low-level forensic information without adapting the pretrained vision encoder. To better evaluate low-level forensic perception, we further construct **SemClean-Bench** through a reverse filtering strategy that removes AIGIs with salient high-level semantic anomalies, shifting the evaluation toward low-level artifacts. Extensive experiments across diverse AIGI benchmarks and SemClean-Bench demonstrate strong and generalizable performance. These results show that strong forensic perception can be achieved by effectively accessing native visual representations without relearning, highlighting representation access as a simple and complementary direction for forensic MLLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.