Pairwise Modality Evidence Transfer for Generalized Zero-Shot Emotion Recognition
Abstract
Generalized zero-shot emotion recognition requires identifying both seen and unseen emotions using task-specific labeled examples only from seen categories. Multimodal observations provide complementary evidence for this task, but audio, text and visual predictors may favor different candidates. Moreover, a correction learned from errors on familiar categories may not remain useful when new categories enter the decision. We propose Pairwise Modality Evidence Transfer (PMET), a shared nonlinear correction of comparative multimodal evidence, to learn a reusable correction from these comparative preferences. We utilize a nonlinear antisymmetric scorer to combine modality margins, the baseline margin and description similarity, generated from semantic recognizer; uniform aggregation converts its pair corrections into residual class scores. To expose the learner to unfamiliar-category predictions during source training, we construct teachers that exclude selected source categories from fitting. Ground-truth query supervision jointly trains uniform and auxiliary relational aggregation, with only the uniform path retained at inference. The formulation changes candidate-level ordering while preserving the underlying semantic recognizer and its calibration policy. Experiments show that our proposed PMET achieves best performance on multiple emotion recognition datasets.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.