SIDE-Bench: Evaluating Social Interaction and Conversational Partner-Intent in Multi-Agent Environments
Abstract
Large language models (LLMs) are increasingly deployed as autonomous agents in complex, multi-user environments where success depends on navigating interpersonal friction. However, existing evaluations primarily test isolated reasoning or cooperative dialogue, failing to capture an agent's ability to manage high-friction social interactions defined by competing goals, hidden constraints, and relational risk. To address this critical gap, we introduce SIDE-Bench, a novel benchmark comprising 150 human-validated, multi-turn scenarios across five domains, designed specifically to test strategic negotiation and social alignment in large language model agents. Because evaluating interaction outcomes alone obscures why an agent fails, we pair SIDE-Bench with the Conversation Partner Intent Tracking (C-PIT) framework. C-PIT tracks an agent's internal belief updates turn-by-turn, allowing us to evaluate whether models can accurately infer a counterpart's intent and ground their subsequent actions accordingly. This combined framework enables researchers to test not only what an agent achieves, but whether its actions remain coherent with its evolving worldview. Evaluating seven frontier and open-weight LLMs reveals a major vulnerability: even state-of-the-art models that achieve perfect internal belief tracking frequently suffer from "action mismatch," generating utterances that violate the very constraints they just correctly inferred. By exposing this critical disconnect between logical reasoning and strategic action-grounding, SIDE-Bench provides a necessary, rigorous tool for developing reliable goal-oriented social agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.