InterRole: Role-Aware Speech Injection for Multi-Character Interaction
Abstract
Audio-driven character animation has achieved impressive progress in single-character scenarios, but remains challenging in multi-character settings. Existing methods often rely on region-exclusive speech injection to prevent identity leakage, which improves speaker-audio binding but limits listener access to speech cues and weakens speaker-listener interactions. In this paper, we identify listener-side behavior as a central component of multi-character audio-driven animation and formulate the key thesis that speech should be globally accessible but decoded in a role-specific manner. Guided by this thesis, we analyze several speech injection schemes and propose InterRole, a complete framework for role-aware multi-character audio-driven animation. Its core module, RoleMixAttn, separates speech conditioning into expressive and receptive attention branches and uses dynamic soft routing to adaptively mix them for each video token. This design jointly models speaker articulation and listener responses while improving robustness to character movement and occlusion. To support training and evaluation, we construct the RoleTalk-AV dataset, adopt a staged training strategy, and introduce interaction-oriented metrics for gaze alignment and prosody-correlated listener motion. Experiments on Wan2.1 and LTX-2.3 show that InterRole generates more coherent speaker-listener interactions than existing multi-character baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.