acceptodds
Under review as a conference paper at ICLR 2027

Same Words, Different Intent: Hostility-Conditioned Contrastive Learning for Robust Multilingual, Multimodal Hate Speech Detection

Abstract

As digital hostility rapidly shifts from text to regional voice notes and multimedia broadcasts, automated moderation systems are failing to keep pace. Deprived of true semantic understanding, current multimodal classifiers act as brittle keyword spotters. They routinely misclassify political sarcasm, non-hateful distress, and pro-group solidarity as toxic abuse simply because these statements contain shared identity terms. Escaping this shortcut-learning paradigm requires nuanced multimodal training, yet there is a critical scarcity of such resources for diverse, low-resource languages. To break this bottleneck, we introduce a massive synthetic hate speech corpus spanning English and nine Indian languages, uniquely aligned across three modalities: text, speech, and text-in-image. Leveraging this foundational resource, we propose a novel hostility-conditioned composite contrastive objective over SONAR and SeamlessM4T encoders. By structuring the latent space around fine-grained hostility categories and integrating minimal-edit counterfactual pairs, we compel the model to learn the true intent behind the language, differentiating genuine malice from benign linguistic overlap. The empirical gains are substantial. On a demanding diagnostic probe isolating these exact linguistic failure modes, our intervention increases accuracy from 78.4% to 95.1%. Furthermore, our approach outperforms the state-of-the-art multimodal baseline across all ten languages. Crucially, this robust representation transfers zero-shot to independent, real-world benchmarks (X-MuTeST and ADIMA). This meaningfully bridges the gap between synthetic training environments and real-world deployment, offering a potential approach for multilingual AI safety.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.