acceptodds
Under review as a conference paper at ICLR 2027

SocialBench : Evaluating Multimodal Perception, Understanding and Reasoning in Social Settings

Abstract

Social understanding requires perceiving behavioral cues, interpreting them in context, and reasoning about implicit intentions. Evaluating only final answers offers an incomplete view of these complementary skills. We introduce SOCIALBENCH, a unified benchmark organizing 16 existing datasets into 30 task variants across three levels: Perception, Understanding, and Theory of Mind. It combines direct assessments of social cues with contextual and inferential tasks across image, video, audio, and text inputs. Evaluation of five multimodal large language models reveals strengths in selected semantic categorization tasks alongside difficulties in fine-grained spatial and temporal prediction and implicit social inference. On Qwen3-Omni, one epoch of task-specific supervised fine-tuning (SFT) substantially improves several initially weak tasks, while joint training brings further gains at the cost of regressions elsewhere. To examine how these improvements relate to internal representations, we train supervised linear predictors on activations extracted before and after SFT. These predictors outperform generated answers on several social tasks, revealing predictive information that the model’s responses do not fully reflect. After SFT, generated answers often improve more than probe predictions, narrowing but not consistently eliminating this gap. Thus, neither zero-shot nor fine-tuned responses alone fully characterize the social competencies evaluated here. SOCIALBENCH provides a framework for studying social capabilities through generated performance, supervised adaptation, and linear decodability together.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.