E-MMSI: Toward Evidence-Grounded Social Interaction Understanding
Abstract
We introduce Evidence-grounded Multi-party Multi-modal Social Interaction (E-MMSI) understanding, a new task for spoken social videos. Beyond answering an interaction question, a model must provide observable evidence that supports its answer, making the prediction verifiable. To study this problem systematically, we draw on interaction theory to build an evidence benchmark on existing social-interaction datasets. The annotation schema represents answer-relevant actors, activities, partners, relations, contexts, and temporal anchors. Evaluation shows that current models often produce unreliable evidence in complex and dynamic multi-party interactions, such as grounding a pointing gesture to the wrong participant or time interval. To address this challenge, we construct a word-level social representation (WSR) that captures dynamic verbal and non-verbal social cues in word interval. Averaged across the three benchmark subsets, WSR improves answer accuracy over the baseline by 3.47 points and evidence score by 13.01 points.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.