acceptodds
Under review as a conference paper at ICLR 2027

Learning to Reason About Implicit Multimodal Harm: From Diagnosis to Generalization

Abstract

A benign image paired with a benign caption can become harmful only in combination, making detection fundamentally a cross-modal reasoning problem. Existing multimodal harm benchmarks evaluate only classification accuracy, leaving unclear whether models identify the cross-modal mechanism that produces harm. We introduce (ltimodal ragmatic arm nterpretation), a controlled diagnostic benchmark in which harm arises purely from image-text composition, with benign counterfactuals and annotated harm rationales. Using MuPHI, we systematically evaluate where compositional harm reasoning breaks in vision-language models (VLMs). MuPHI reveals that VLMs reliably ground each modality yet fail to bind the evidence into the harm mechanism and their predictions respond weakly to label-determining counterfactual interventions. Moreover, existing harm detection approaches trained on labels substantially improves in-domain (ID) detection but transfers poorly across harm distributions. We therefore propose ompositional arm nterpretation earning (), which rewards evidence identification in each modality, cross-modal binding and verdict consistency in the reasoning chain rather than just predicted label. CHIL improves 81% OOD transfer evaluations by an average of macro-F1 and correct counterfactual flip rate by up to and % points, while also improving overall reasoning quality over label-only optimization. These results suggest that optimizing how a model reasons, not just what it concludes, offers a promising direction for multimodal safety systems that generalize beyond dataset-specific correlations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.