OVGER: Evidence-Grounded Open-Vocabulary Group Emotion Recognition via Blind Annotation and Progressive Reward Injection
Abstract
Group Emotion Recognition (GER) infers collective affect from multi-person audio-visual recordings for social-interaction analysis, multimedia understanding, and human-centered computing. Most existing GER systems formulate the task as closed-set classification over coarse affect categories. Although effective for benchmark comparison, this discriminative interface provides limited fine-grained affect description and little explicit evidence for the prediction. Motivated by this gap, we formulate open-vocabulary GER as a benchmark-compatible extension beyond conventional closed-set prediction and instantiate it with OVGER. The model retains the coarse group-affect decision while adding more specific free-form emotion descriptors and an evidence-linked structured response. The central contribution is the supervision-and-reward interface that makes these outputs jointly trainable and auditable: OV-VGAF augments the Video-level Group AFfect (VGAF) benchmark with label-blind modality observations, structured supervision, and hidden evidence references for evaluating coarse labels together with open-vocabulary affect and evidence alignment. This design couples benchmark-compatible group decisions with fine-grained affect descriptions whose supporting multimodal cues can be inspected. Extensive experiments show competitive coarse-label performance relative to representative task-specific GER systems and clear gains over the evaluated zero-shot MLLMs on the same 766-clip validation split, reaching 74.67% accuracy, 74.50% macro-F1, and 85.51% open-vocabulary emotion soft-F1.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.