acceptodds
Under review as a conference paper at ICLR 2027

Embodied Interaction Policy: A Unified Model for Real-Time Human-Robot Interaction

Abstract

Real-time human-robot social interaction requires continuous coordination between speech and physical behavior. Existing systems often separate interaction agents from motion generation, complicating coordination as conversational states change. We propose Embodied Interaction Policy (EIP), a unified streaming framework that couples an Audio-Interaction expert with a motion expert through layer-wise joint attention. This allows motion generation to directly access the evolving interaction context. A chunk-wise causal mask prevents access to future context, while a vocal token encodes speaking state and motion intensity derived from the response emotion. We construct Embodied Interaction Dataset (EI-Data) from co-speech and dyadic interaction data and train our model in two stages. EIP supports proactive responses to salient acoustic events, state adaptation between listening and speaking, and interruption-aware turn-taking. Quantitative evaluations demonstrate improved motion quality over representative gesture-generation baselines. Compared with RoboGesture, EIP achieves higher interaction success rates and reduces warm-start latency from speech input to the first generated motion chunk from 3.78 s to 1.35 s. Physical deployment and user studies further demonstrate improved motion naturalness and interaction comfort.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.