acceptodds
Under review as a conference paper at ICLR 2027

Can You Find the Edit When Nothing Looks Wrong? A Paired Benchmark for Manipulation Localization Beyond Semantic Abnormality

Abstract

Instruction-driven image editing can produce manipulations that remain semantically coherent with the surrounding scene. Standard per-image evaluations do not reveal whether a localizer detects forensic evidence or relies on semantic abnormality that coincides with the edited region. We formalize this issue as the semantic-plausibility shortcut: a model may localize an anomalous replacement yet miss a plausible edit of the same target. We introduce Plausible–Anomalous Paired Edits (PAPE), a benchmark that makes intended semantic plausibility an explicit evaluation factor. Each triplet contains a source image and two edits of the same target, one plausible and one anomalous, produced with one editor and matched controlled editing settings, independently annotated masks, and overlap filtering for comparable spatial support. A two-stage construction pipeline yields ∼47K bilingual triplets from 6.5K photographs shared across four editors, three source domains, and six semantic violation categories. A pair-preserving protocol reports localization quality in each condition and the gap between them. Across seven public localizers, editor-averaged IoU is lower on plausible edits for every method on ImageNet and Places and for six of seven on COCO, with positive relative gaps from 2.5% to 42.2%. PAPE exposes a reliability dimension that aggregate scores conceal and provides paired records for analyzing its causes. Our anonymized code and dataset is available at https://anonymous.4open.science/r/PAPEDataset/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.