Meaning Beyond Language: Learning Complementary Semantics for Multimodal Metaphor Detection
Abstract
Metaphor is widely used to express meanings beyond literal content. Multimodal metaphor detection aims to identify metaphorical meaning by jointly understanding text and images. Although existing methods have made notable progress, they mainly rely on a single language and overlook the complementary semantic information expressed in other languages. To address this limitation, we propose \ourmethod, a framework that uses an auxiliary language to improve multimodal metaphor detection. First, we introduce a cross lingual semantic disentanglement module to separate shared semantics from language specific information. Contrastive learning, reconstruction, and orthogonality constraints are used to preserve useful information and reduce the effect of language differences. Second, since different vision language model layers capture different semantic cues, we fuse the representations of both languages at each layer through self attention. We then assign different weights to these layers based on the input sample, allowing the model to focus on the semantic cues most relevant to metaphor detection. Finally, to address the absence of translations during inference, we learn a mapping that predicts the semantic representation of the translation from that of the original text. Experiments on multiple benchmarks with different vision language models demonstrate the effectiveness of our framework.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.