acceptodds
Under review as a conference paper at ICLR 2027

MMSec-R1: Test-Time Scaling for Fake Image Reasoning

Abstract

Current Multimodal Large Language Models (MLLMs) typically treat image forgery detection and localization as separate or loosely coupled tasks, often relying on massive supervised data. In this paper, we redefine this paradigm by proposing Fake Image Reasoning (FIR), a novel task that equates authenticity verification with the rigorous reasoning of forgery traces—essentially, an image is fake only if specific manipulated regions can be reasoned out. We introduce MMSec-R1, the first framework to explore Test-Time Scaling (TTS) for FIR via a multi-stage unified Reinforcement Fine-Tuning (RFT) strategy. Unlike traditional methods, MMSec-R1 treats precise localization as the definitive correctness verification, compelling the model to self-evolve thinking patterns from limited cold-start data. We validate our approach by training models across different scales (4B and 8B), both achieving superior performance. Our results demonstrate that scaling inference-time compute via RL effectively induces complex forensic reasoning in compact models, bridging the gap between detection confidence and visual evidence.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.