acceptodds
Under review as a conference paper at ICLR 2027

SocialAgent: From Collaborative Reasoning to Self-Improving Social Intelligence

Abstract

Artificial Social Intelligence (ASI) refers to the ability of machines to understand and respond to human social behavior, constituting an important dimension of general artificial intelligence. Such ability depends on understanding how different participants perceive the same situation and drawing on experience from previous interactions. However, current multimodal large language models (MLLMs) largely reason about each interaction in isolation, without explicitly reconciling participants' perspectives or accumulating reusable social knowledge. We introduce SocialAgent, a framework that develops social reasoning through collaborative experience acquisition and its subsequent internalization into a single model. SocialAgent-Scaffold combines Role-based Collaboration (RoC) with an Evolving Experience Pool (EEP) without parameter updates. RoC reasons from individual participants' perspectives and reconciles their complementary evidence, producing explicit demonstrations of perspective taking. The EEP abstracts successful reasoning processes into reusable role and reasoning templates, turning individual solutions into strategies that can guide new interactions. Although this collaboration provides structured reasoning demonstrations, repeatedly simulating multiple roles incurs substantial inference cost. SocialAgent-Evolve therefore internalizes this experience into a single-agent MLLM through Strategy-Mediated Self-Refinement (SMSR). Its Strategy-guided Learning stage trains the student on scaffold-generated trajectories paired with their guiding EEP strategies. The trajectories demonstrate how to solve concrete cases, while the strategies make generalizable reasoning principles explicit. This pairing enables the student to learn perspective taking without replaying multi-role collaboration at inference time. Since imitation alone leaves gaps beyond the demonstrated solutions, Multi-strategy Policy Optimization then targets the student's remaining failures, using alternative EEP strategies to guide exploration of diverse reasoning paths. Together, these stages connect learning from collaborative demonstrations with failure-driven self-refinement, without relying on a stronger external teacher. Extensive experiments on social reasoning benchmarks demonstrate that SocialAgent achieves state-of-the-art performance on both open-ended generation and multiple-choice question answering.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.