acceptodds
Under review as a conference paper at ICLR 2027

E-MMSI: Toward Evidence-Grounded Social Interaction Understanding

Abstract

We introduce Evidence-grounded Multi-party Multi-modal Social Interaction (E-MMSI) understanding, a new task for spoken social videos. Beyond answering an interaction question, a model must provide observable evidence that supports its answer, making the prediction verifiable. To study this problem systematically, we draw on interaction theory to build an evidence benchmark on existing social-interaction datasets. The annotation schema represents answer-relevant actors, activities, partners, relations, contexts, and temporal anchors. Evaluation shows that current models often produce unreliable evidence in complex and dynamic multi-party interactions, such as grounding a pointing gesture to the wrong participant or time interval. To address this challenge, we construct a word-level social representation (WSR) that captures dynamic verbal and non-verbal social cues in word interval. Averaged across the three benchmark subsets, WSR improves answer accuracy over the baseline by 3.47 points and evidence score by 13.01 points.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.