acceptodds
Under review as a conference paper at ICLR 2027

AIRIS: Benchmarking Social Embodied Intelligence in Egocentric Streaming Videos

Abstract

Embodied agents operating around people must not only recognize social cues, but also decide when the available evidence is sufficient, whether to intervene, and how to respond without violating social norms. Existing benchmarks typically evaluate these abilities in isolation or with access to completed videos, obscuring where an online social decision fails. We introduce AIRIS, a causal streaming benchmark that formulates social embodied intelligence as a time-anchored chain of Attention, Intention, Reaction, and Interaction (A-I-R-I). AIRIS contains 8,000 expert-curated questions grounded in 357 egocentric clips across 11 context categories and 36 capability dimensions. At each query, models receive only the observation prefix available at that time; decision-aware metrics reward calibrated waiting and complete reactions while penalizing premature commitment. Evaluating a pool of 45 proprietary and open-source general-purpose, embodied, streaming-video, and omni-modal models reveals fragmented and temporally miscalibrated competence. General-purpose VLMs perform strongly in cue grounding and intent recognition, whereas embodied specialists are more competitive on aggregate Reaction; nevertheless, the best Intent Anticipation Score is only 0.237, category-level False Commitment Rates range from 0.26 to 0.51, and no model exceeds 0.147 Reaction Joint-EM. These findings characterize the perception–execution gap as a failure of temporal calibration and reaction composition. Ultimately, reliable social embodied intelligence requires models to actively calibrate anticipation to observable evidence, composing perception, inference, and action into seamless responses that align with human social norms.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.