TD-AFGW: TONE-AWARE MULTIMODAL EMOTION TRANSFER WITHOUT TARGET EMOTION LABELS
Abstract
Transferring emotion recognition to tonal languages with scarce annotations requires addressing cross-lingual representation shift and the overlap between lexical tone and emotional prosody. We introduce , a multimodal framework trained with labeled source conversations and target conversations without emotion labels. A text-conditioned pitch predictor, ToneNet, estimates a lexical-tone baseline and supplies residual pitch features and feature-wise acoustic conditioning. Anchored fused Gromov–Wasserstein alignment combines within-domain relations with fixed facial and paralinguistic cues. The transport plan propagates source labels, while a tone-swap consistency gate selects target soft labels for distillation. A visual-to-audio-to-text curriculum progressively enables adaptation. We study English-to-Mandarin transfer from MELD to EmotionTalk, using Mandarin as an established tonal-language benchmark for task-specific annotation scarcity. Under shared source-prior logit calibration, the full model achieves weighted F1 and macro F1 across three seeds, versus and for source-only transfer. Ablations support the contributions of anchored geometry, ToneNet, and consistency filtering. Per-class results show particularly strong gains for anger and surprise, alongside persistent minority-class weaknesses. The findings support tone-aware transfer on this benchmark; broader low-resource generalization and fully target-label-blind model selection remain to be established.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.