DINOv3 Beats Specialized Detectors: Revisiting Image Manipulation Localization from a Foundation Model Perspective
Abstract
Image manipulation localization (IML) identifies regions that have been spliced, copied, removed, or otherwise edited. Its labels describe provenance rather than semantic identity, so existing systems rely on specialized forensic cues. Frozen DINOv3 features, despite strong readouts on general vision tasks, localize manipulations poorly. We ask whether a general backbone can make editing evidence readable without a dedicated forensic front end. A stronger decoder helps only in part: a global Transformer head improves frozen features, but adapting the encoder yields a far larger gain. We call this contrast the forensic readability gap: frozen features contain editing information, but it is difficult to read with one linear rule shared across images. We therefore propose DINOv3-IML, a simple QKV-LoRA adaptation with a shallow convolutional head and no forensic feature extractor. Adaptation makes editing evidence readable through a shared linear rule and aligns editing directions across images; once the encoder is adapted, a global head brings no further gain. DINOv3-IML reaches 0.85 and 0.77 pixel F1 under the CAT-Net and MVSS protocols, surpassing the best prior methods by 12 and 24 points, respectively, and establishes a simple yet strong baseline for IML.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.