GESBENCH: When and Where Do Multimodal Language Models Perceive Social Meaning?
Abstract
Human social perception is incremental: we infer intent before an utterance is complete, revise our interpretation as new cues arrive, and attend to specific facial, gestural, or prosodic signals that shape social meaning. Current multimodal LLM benchmarks largely evaluate final-answer understanding over complete episodes, rewarding post-hoc recognition of social meaning rather than the unfolding perception required for real-time interaction. Drawing on interactionally embedded Gestalt principles of multimodal human communication, we introduce GesBench: a human-grounded benchmark for cue-level social perception in multimodal LLMs. GesBench uses cumulative sub-utterance clips to measure when models form stable interpretations of speaker intention and social affordance, and human gaze-guided corruption to evaluate whether regions attended by humans are functionally important for model judgments. Built from more than utterance-level clips from scripted and unscripted conversational datasets, GesBench includes human annotations of unfolding judgments and gaze on a representative subset. Across open and closed multimodal models, we find a clear human-model gap: humans form predictive understandings early and revise more often, while models often hedge or repeat across different clips. GesBench shifts social evaluation from whether models eventually recognize meaning to when and where that meaning is formed. Our codes are publicly available at https://anonymous.4open.science/status/GesBench-share-212C.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.