A Progressive Multi-Stage Multimodal Dataset and Benchmark for Real-World Chinese Emotion Recognition
Abstract
Real-world emotion datasets must jointly balance acquisition control, ecological validity, cross-modal synchronization, label reliability, and reproducible evaluation. We construct a progressive multi-stage multimodal dataset for Chinese emotion research with 68,210 speech segments from 50 participants, nine perceived-emotion labels, synchronized timestamped photo sequences, heart-rate-variability (HRV) measurements, and inertial signals. A unified protocol links individual emotion elicitation, dyadic interaction, and daily-life recording while retaining the same participant cohort, label space, temporal reference, and quality controls. To address the subjectivity of perceived emotion, an auditable annotation pipeline combines independent ratings, absolute-majority aggregation, and adaptive annotator expansion. Across 65,456 segments with at least two alignable ratings, Fleiss' and Krippendorff's ; random resampling of 517 samples with ten complete ratings and fixed-model learnability tests further show that label agreement and learnability generally improve with more annotators. We establish unified, speaker-independent seven- and nine-class benchmarks; the complete seven-class benchmark covers 13 speech baselines, with matched-model comparisons on MELD and IEMOCAP. Chinese-HuBERT-Large achieves 78.42% Accuracy and 52.59% Macro-F1 on the seven-class task. A cumulative three-stage curriculum raises UAR from 51.63% to 54.07%, while yielding 76.61% Accuracy and 52.03% Macro-F1. The strict complete-modality intersection contains 37,043 segments (54.31%). On this subset, additional visual and sensor modalities mainly produce class-conditional performance changes: visual input improves accuracy for six non-neutral classes, while overall Macro-F1 remains close to the audio baseline. Together, the dataset, progressive acquisition framework, empirically validated consensus protocol, and systematic benchmark provide an auditable foundation for real-world Chinese emotion research.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.