acceptodds
Under review as a conference paper at ICLR 2027

TraceTok: Fine-Grained Text Tampering Localization via Native Mask Token Generation

Abstract

Text images are pervasive in financial, legal, and administrative records, where text tampering can undermine the credibility of the underlying information. Existing forensic multimodal large language models (MLLMs) can localize suspicious regions and provide natural-language analysis. However, they localize tampering through textual bounding boxes or tokens coupled to segmentation decoders, which are coarse for irregular shapes and fragmented regions not aligned with object-level semantic boundaries. Moreover, MLLM visual encoders are less sensitive to subtle manipulation traces than dedicated forensic experts. To address these limitations, we propose TraceTok, a forensic-enhanced MLLM that formulates fine-grained tampered text localization as native mask-token generation and jointly produces textual forensic responses and pixel-level tampering masks. The Tamper-Aware Adaptive-Length Mask Reconstruction module (TAMR) conditions mask reconstruction on image-side forensic features and uses Token Length Predictor (TLP)-guided retained-length selection, enabling compact token prefixes to preserve small and irregular tampered text structures. Evidence-Preserving Adaptive Grid Fusion (EAGF) aligns multistage expert features with the current visual-token grid before fusing and injecting them into the visual branch, thereby strengthening low-level forensic perception across varied input resolutions. On TFR-Test, CIS, and CTM, TraceTok attains 62.2 average Mask IoU, exceeding our reevaluated TextShield-R1 baseline by 12.7 points; on RTM it matches the strongest task-specific localizers, and its reconstructor reaches comparable or better F1 with fewer retained tokens than previous methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.