acceptodds
Under review as a conference paper at ICLR 2027

Do Forensic Verdicts Follow Output Regions?

Abstract

Explainable forensic VLMs pair manipulation verdicts with localized evidence, but localization overlap alone cannot establish whether the verdict actually depends on the cited region. Existing deletion and inpainting tests are also problematic because the intervention can itself introduce forensic cues. We introduce paste-back auditing: when a paired manipulation retains its authentic source, we restore authentic pixels only inside the model's output region and measure the resulting verdict change (necessity) alongside localization. Paste-back introduces no synthesized content and is exactly the identity on authentic images, whereas two classical inpainters flip 31.3% of authentic examples for served chat VLMs. In a positive control, a faithful oracle and a correct-but-non-causal oracle have identical IoU, yet necessity is 0.99 versus 0.00, showing that localization and intervention response capture distinct properties. Controlled Qwen2-VL-2B experiments reproduce this dissociation under a JPEG-75 cue. We further audit FakeShield and FFAA as released-system case studies. The audits reveal that output regions can be only partially aligned with the measured verdict response, while narration- grounded interventions can produce stronger effects when the narration explicitly identifies the manipulated region. Paste-back therefore provides a synthesis-free test of verdict–output-region alignment wherever the authentic source is retained, while keeping localization correctness and intervention dependence as separate quantities.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.