SAFE-Align: Alignment-Conditioned Evidence Routing for Multimodal Manipulation Detection and Grounding
Abstract
Multimodal manipulation detection requires reasoning over image–text pairs whose semantic correspondence can vary substantially across samples. Existing methods typically use multimodal fusion mechanisms whose aggregation is not explicitly conditioned on the sample-level image–text alignment regime. We challenge this assumption and propose SAFE-Align, an alignment-conditioned framework that treats image–text correspondence as a sample-level routing prior rather than an authenticity score. SAFE-Align combines a discrete alignment state, obtained by quantizing global image–text similarity and represented by a learned state embedding, with the original continuous similarity to characterize the correspondence regime, and uses this alignment prior to route complementary global-interaction, local-spatial, and text-only evidence. To preserve fine-grained manipulation cues, it further models bidirectional patch–token correspondence and captures local agreement and discrepancy cues for image and text grounding. On the official DGM4 split, SAFE-Align achieves 95.63% AUC, 88.74% accuracy, and 89.63% mAP, improving over ASAP by 1.25, 1.03, and 1.10 percentage points, respectively. Alignment-conditioned routing also improves AUC by 1.55 points over uniform fusion and by 1.97 points over a learned gate without the alignment prior. These results support explicitly conditioning multimodal evidence routing on cross-modal alignment rather than relying on a fixed fusion pattern.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.