Learning Transferable Anti-Forensic Priors via Adversarial Diffusion Training
Abstract
Image Manipulation Detection and Localization (IMDL) models have achieved remarkable progress in identifying and localizing manipulated regions, yet their robustness against anti-forensic attacks remains largely unexplored. Existing adversarial approaches typically rely on victim-specific information or image-by-image optimization, limiting their scalability and practical applicability. Inspired by recent advances that repurpose pretrained diffusion models as visual priors for diverse downstream tasks, we investigate whether such priors can be adversarially activated to acquire generalizable anti-forensic capability. To this end, we propose A-PAIR (Adversarial Prior Activation and Image Refinement), a framework that learns a transferable anti-forensic prior through adversarial adaptation of pretrained diffusion representations. Rather than optimizing perturbations independently for each image, A-PAIR trains a diffusion-based attack branch in competition with a manipulation localization branch. The two branches are optimized alternately: the localization branch learns to recover manipulation masks from generated adversarial images, while the attack branch learns to generate images that induce incorrect localization. Through this alternating competition, the diffusion representation acquires a transferable anti-forensic prior that can directly generate adversarial images across different inputs and localization models. The refinement stage improves image quality while preserving attack effectiveness, and further strengthens attacks when victim queries or gradients are available. Extensive experiments across multiple IMDL benchmarks and localization models demonstrate the advantages of A-PAIR over existing anti-forensic methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.