Fine-Grained Image Tampering Localization via Semantic Proposals and Adaptive Residual Refinement
Abstract
Proactive image protection enables subsequent tampering to induce observable responses in pretrained vision models, yet these responses are often too coarse and spatially inconsistent to support fine-grained localization. We investigate this observability-to-localization gap in a proactive protection setting, where a protected reference and its manipulated counterpart are available. We propose SPARR (Semantic Proposal and Adaptive Residual Refinement), a response-interpretation framework that converts manipulation sensitive responses into accurate tampering masks while keeping the upstream protection mechanism unchanged. SPARR first exploits paired response variations under shared spatial prompts to derive robust semantic proposals, and then uses paired RGB residuals to refine response-induced regions toward the actual manipulated structures. To account for sample-dependent reliability, SPARR further introduces an adaptive routing strategy that evaluates whether additional refinement is supported by ground-truth-free evidence and selectively accepts or rejects it. This design enables SPARR to improve localization without relying on a fixed refinement behavior across different manipulation patterns. We evaluate SPARR under controlled same-protection comparisons, five conventional image tampering benchmarks, and four generative-AI editing settings. The results show that reliable downstream response interpretation can substantially improve fine-grained localization without modifying the upstream protection mechanism, while maintaining strong performance across diverse editing settings. The code will be made publicly available upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.