Reading the Visual Clues: A Dataset and Contextual Detector for Document Tampering
Abstract
Document tampering extends beyond edited text to figures, formulas, and tables, motivating a common evaluation of editing operations across page elements. We introduce MEDTD, a benchmark for localizing operation-involved regions and recognizing editing operations from a single full-page image. Constructed from DocLayNet, MEDTD contains 46,290 programmatically edited pages, 25,145 untampered counterparts, and 69,839 region annotations covering seven element groups and four operations. Its annotations include unchanged duplication sources and vacated relocation sources, rather than modified pixels alone. We further introduce ClueDet, a DINO-based detector that refines detection queries through similarity-weighted comparisons of predicted regions. Under matched training settings, ClueDet achieves 34.11 operation AP, compared with 33.02 for DINO and 32.56 for a parameter-matched uniform-context variant. Improvements are concentrated in textual elements and splicing rather than shared uniformly across element groups. Evaluation on untampered pages and external datasets characterizes false-positive behavior and limits to transfer. MEDTD provides a benchmark for studying operation-aware document tampering across heterogeneous page elements.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.