acceptodds
Under review as a conference paper at ICLR 2027

Detecting Implicit Sarcasm in Speech by Cross-modal Reasoning with Multi-source Knowledge

Abstract

This paper focuses on detecting complex sarcasm emotions in speech. Simple emotions like happiness can be directly judged based on acoustic signals in speech. By contrast, sarcasm involves a conflict between surface features and underlying intention, e.g., using positive words to convey sadness. That is hard to detect by the sound or words alone. Moreover, implicit sarcasm often involves external knowledge related to commonsense, plots, scenes, etc., which is not presented in the given speech, resulting in insufficient clues for inferring the conflict. To solve these issues, we propose a new reasonable framework. We first encode the text and acoustic features in speech, and grasp their subtle correlations. That helps to find potential sarcastic conflicts in the heterogeneous space. We then retrieve the missing textual and prosodic knowledge related to the speech. They are necessary clues for inferring conflicts and are used to build an emotional graph. Based on it, we deduce a conflict chain between the literal words, true intention, and knowledge constraints by reinforced multi-hop inference. This chain can help to identify implicit sarcasm more accurately and explainably. Moreover, we create a large-scale dataset called AEbSD to conduct extensive evaluations. The results show our effectiveness and reliability.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.