acceptodds
Under review as a conference paper at ICLR 2027

ForRAIL: Forensic Retrieval-Augmented Visual In-Context Learning for Training-Free Image Forgery Localization

Abstract

Image forgery localization (IFL) has become increasingly challenging as modern image editing models produce realistic and semantically coherent manipulations, while existing methods often rely on task-specific training and struggle to generalize to unseen manipulation techniques and generators. Multimodal large language models (MLLMs) offer broad visual reasoning capabilities and cross-domain knowledge, yet their potential for precise, training-free forgery localization remains largely unexplored. In this work, we investigate whether a frozen MLLM can perform pixel-level IFL through visual in-context learning. We propose Retrieval-Augmented Visual In-Context Learning for IFL (ForRAIL), a training-free framework that retrieves informative forgery demonstrations for each query to guide localization without updating model parameters. To improve demonstration selection, we further introduce CoSFoR, a coarse-to-fine semantic–forgery retrieval strategy that jointly considers semantic relevance and manipulation-sensitive characteristics. Extensive experiments across diverse out-of-domain datasets, generators, and multiple MLLM backbones show that ForRAIL substantially outperforms vanilla MLLMs and existing training-free MLLM-based approaches and exceeds the best specialized trained forensic models 0.03 IoU on OOD average. It also maintains strong localization performance across different tampered-region sizes and under common image perturbations, including Gaussian blur, Gaussian noise, JPEG compression, and scale. These results establish retrieval-augmented visual in-context learning as a promising training-free paradigm for generalizable image forgery localization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.