MMSec-R1: Test-Time Scaling for Fake Image Reasoning
Abstract
Current Multimodal Large Language Models (MLLMs) typically treat image forgery detection and localization as separate or loosely coupled tasks, often relying on massive supervised data. In this paper, we redefine this paradigm by proposing Fake Image Reasoning (FIR), a novel task that equates authenticity verification with the rigorous reasoning of forgery traces—essentially, an image is fake only if specific manipulated regions can be reasoned out. We introduce MMSec-R1, the first framework to explore Test-Time Scaling (TTS) for FIR via a multi-stage unified Reinforcement Fine-Tuning (RFT) strategy. Unlike traditional methods, MMSec-R1 treats precise localization as the definitive correctness verification, compelling the model to self-evolve thinking patterns from limited cold-start data. We validate our approach by training models across different scales (4B and 8B), both achieving superior performance. Our results demonstrate that scaling inference-time compute via RL effectively induces complex forensic reasoning in compact models, bridging the gap between detection confidence and visual evidence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.