Thinking with Spatial–Social Code for Social Reasoning in Videos
Abstract
Real-world social reasoning is essential for intelligent systems to understand human behavior, anticipate needs, and coordinate actions with people. Inferring beliefs, goals, and intentions requires connecting the visual context of an interaction with each participant's history of observations and actions. The central challenge is bridging observable interaction and latent social states: visual events acquire social meaning through a participant's access to information, prior experience, and ongoing goals. A representation must therefore capture not only the scene, but also how it informs different participants' perspectives. We argue that explicit code offers a computational medium for organizing and using these dependencies through stable references, agent-specific records, and update operations. We introduce Spatial-Social Code (SSC), which links physical situations, visual evidence, and inferred social states in a persistent, updateable representation. Spatial Code organizes entities, spatial relations, physical events, and perceptual-access evidence. Social Code links these records to participants' information and hypotheses about their beliefs, goals, and intentions. Their connections make explicit how situated experiences inform social inference, supporting both queryable reasoning and social memory over time. We evaluate SSC on cross-agent video question answering, collaborative visual grounding, and embodied interaction, demonstrating improvements across multiple tasks. Together, these capabilities provide a computational interface for social reasoning grounded in evolving physical situations and individual experience.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.