Awakening Verbalizer Forensics: Training-Free Readout Alignment for AI-Generated Image Detection in MLLMs
Abstract
Multimodal large language models (MLLMs) show promise for AI-generated image detection, but existing approaches mainly fine-tune them with binary real/fake supervision that, even with explanatory objectives, easily overfits to training fake patterns and fails to generalize. We find that frozen MLLM representations already contain transferable forensic information that the pretrained verbalizer, i.e., vocabulary projection, does not fully exploit. We introduce Forensic Readout Alignment (FRA), which estimates a forensic direction with a boundary offset using real images and single generator's fake images, and encodes them into the vocabulary projection weights of the real/fake answer tokens, without updating the backbone or any gradient-based training. FRA preserves the original distance and midpoint between the two answer rows, keeping the fitted boundary calibrated against the rest of the vocabulary. Using the same training data as the evaluated methods, FRA achieves the best overall cross-generator detection performance, outperforming LoRA-based adaptations, and remains effective across different training generators, training-set sizes, and MLLM backbones. Overall, FRA provides a lightweight way to effectively awaken pretrained forensic capabilities, while preserving the backbone's general language and visual understanding capabilities.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.