Know What You Don't Know: Confidence-Calibrated MLLMs for Tampering Text Detection
Abstract
Multimodal large language models (MLLMs) can detect, localize, and explain text-image tampering, but lack explicit confidence outputs, which weakens the reliability of detection systems in practical scenarios. While calibration methods allow models to output confidence matching prediction accuracy, they optimize confidence separately and only consider correctness in calibration. Under such paradigm, models cannot learn forensically beneficial knowledge from calibration, nor can they perceive the certainty of their own analytical process. To address this, we propose a three-stage confidence-calibrated training framework that connects self-assessment with forensic learning. First, forensic continual pretraining on large-scale text-image data develops task-specific perceptual knowledge. Second, supervised fine-tuning introduces structured reasoning, answer, and confidence outputs together with a **Confidence Grounding Attention Module** (CGAM). CGAM uses the confidence representation to aggregate selected reasoning and answer representations, with auxiliary authenticity supervision tying this pathway to forensic prediction. Third, Group Relative Policy Optimization jointly optimizes forensic rewards and a **Forensic-Quality-Aware Calibration** (FQAC) reward. The calibration target combines current-rollout correctness and cross-rollout accuracy with token-level reasoning certainty, using the latter as an auxiliary process signal rather than a substitute for empirical correctness. Experiments on four test sets show that our framework outperforms TextShield-R1 and achieves an average Expected Calibration Error of 21.4%, lower than the evaluated calibration baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.