acceptodds
Under review as a conference paper at ICLR 2027

Region-to-Fake: Perceptual Reinforcement learning for Fine-Grained Forgery Reasoning

Abstract

While Multimodal Large Language Models (MLLMs) significantly enhance forgery perceptual image understanding, they still exhibit limited perceptual reasoning capabilities when applied to the specialized domain of fine-grained image forgery detection. To address these issues, we propose the Region-to-Fake Distillation (R2F) framework. Our approach leverages large-scale MLLMs combined with image-cropping tool invocation to distill multi-round image crops and bounding box coordinates directly into the Chain-of-Thought (CoT). This enables the student model to acquire internalized bounding box localization capabilities, effectively eliminating the reliance on external tools during inference. To further optimize the model's perceptual understanding of forgeries while mitigating the prolonged inference latency, we introduce Perceptual-Aware Reinforcement Learning with Verifiable Rewards (PA-RLVR). By leveraging the intrinsic tag-generation capabilities acquired during the cold-start phase, this strategy enables direct importance sampling at the tag-sequence level, which significantly enhances the performance of forgery localization. Extensive experiments demonstrate that our method achieves high performance and efficient inference latency across multiple Fake Image Detection and Localization benchmarks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.