acceptodds
Under review as a conference paper at ICLR 2027

Evolving Multi-modal Large Language Model Without Training for Zero-shot Image Forgery Localization

Abstract

Image Forgery Localization(IFL) aims to identify and precisely locate tampered regions in images. To tackle this practical but challenging task, existing methods typically detect forged region in different scenarios by employing a large amount of pre-collected training data. However, with the rapid iteration and upgrading of forgery strategies, new forgery localization requirements are often accompanied by training data that cannot be collected in advance, making it difficult for existing methods to fail against the latest forgery techniques. In this article, we focus on a such practical and challenging task, namely Zero-shot Image Forgery Localization(ZIFL). Inspired by recent advances in Multi-modal Large Language Models(MLLMs), we propose a training-free image forgery localization method called Collaborative LEAder-worker discrimiNator, called CLEAN, which performs zero-shot forgery localization by gathering the cross-modal reasoning capacity of MLLMs and the fine-grained localization capacity of Segment Anything(SAM) to identify forgery clues. Specifically, we propose a Coarse-Grained Suspect Exploration(CGSE) module to act as a forensic analysis leader and systematically generate a list of candidate suspect regions based on high-level forgery evidence. Besides, a Fine-Grained Iterative Refinement(FGIR) module is designed to judge the forgery clues in an iterative manner, which eliminates normal backgrounds and refines the boundaries of tampered regions for accurate ZIFL. Extensive experiments on multiple benchmarks demonstrate that our method achieves competitive IFL performance against existing training-based methods without relying on pre-collected training data on zero-shot image forgery localization task. Our code will be released soon.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.