RoGRE: Image Manipulation Localization as Controlled Hypothesis Revision
Abstract
Image manipulation localization (IML) requires combining heterogeneous evidence with different localization capabilities and failure modes. Existing methods typically fuse such representations as equal-status features, obscuring how each source should modify the current prediction. We instead formulate IML as controlled hypothesis revision: localization begins from an initial estimate, and complementary evidence progressively corrects its remaining errors. Based on this view, we introduce RoGRE, which maintains a shared logit-valued localization state and revises it through role-specific stages. DINOv3 first establishes an initial hypothesis, SAM2 features correct its spatial organization, and forensic residual features refine manipulation evidence that remains weakly represented. Each refinement proposes a bounded signed correction whose magnitude is regulated by a gate conditioned on both the incoming state and stage-specific evidence, allowing informative revisions while limiting unreliable updates. The evolved state is finally decoded by SAM2 into the manipulation mask. RoGRE achieves mean pixel-level F1 scores of 0.888 and 0.813 under the CAT-Net and MVSS-Net protocols, respectively. Ablations further validate the correction order and state-dependent gating, supporting controlled hypothesis revision as an effective approach to heterogeneous evidence integration for IML.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.