Detect the Forgery, Not the Visual Disturbances: MLLM-Guided Representation and Selective Modulation for ID Document Images
Abstract
ID document image tampering detection and localization is an important problem in multimedia security, where detection determines whether an image has been tampered with, while localization identifies the tampered regions. Existing methods primarily focus on clean ID document image conditions. However, Visual Disturbances (VD), such as glare, mosaics, or stains, may occur in ID document images; their low-level features are highly similar to tampered artifacts, potentially leading to VD Induced Forensic Misattribution (VDFM). To address this problem, we propose CLEAR-ID, a new approach for ID document VD scenarios that integrates forensic experts with Multimodal Large Language Model (MLLM). CLEAR-ID uses prior information introduced by an MLLM to adaptively represent forensic expert evidence as VD-oriented or tampered-oriented cues, and projects the resulting features into a low-rank residual space to constrain the modulation of forensic expert evidence. To fill the data gap in this scenario, we further introduce DISTURBID-180k. It covers 42 types of VD and provides corresponding severity and localization mask annotations, enabling systematic analysis of the impact of VD on tampered detection and localization. Experiments demonstrate that VD-aware evidence representation and constrained VD-guided modulation effectively mitigate VDFM. Across diverse VD, CLEAR-ID improves both image-level detection and pixel-level localization over the state-of-the-art (SOTA) forensic expert baseline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.