acceptodds
Under review as a conference paper at ICLR 2027

SocialWorldBench: Benchmarking Social Understanding and Prediction

Abstract

Can AI understand a conversation and predict how it will end? We introduce SocialWorldBench, a benchmark of 19,107 questions drawn from six existing datasets. Its main test uses 396 human negotiation dialogues. For each dialogue, a model reads the same conversation excerpt and answers two questions: what negotiation strategies are being used, and how satisfied will a participant be afterward? We compare the satisfaction prediction with what the participant later reported. This lets researchers examine whether models that recognize negotiation strategies also make accurate predictions. We test five large language models from one provider on a 1,401-question sample and the paired negotiation task. When scored by answer-label overlap (Jaccard), three models identify strategies worse than a simple method that always gives the most common answer. Changing the evaluation questions also reverses two models’ rankings, even though their answers stay unchanged. SocialWorldBench provides researchers with an evaluation dataset to test whether AI can understand conversations and predict their outcomes.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.