acceptodds
Under review as a conference paper at ICLR 2027

HHI-Dyn: Learning Interaction Dynamics for Human-Human Motion Diffusion

Abstract

Text-driven human interaction generation demands a robust latent space structure, where each individual's motion patterns are jointly shaped by the textual description and the partner's behavior. Existing interaction models neglect the coupling between local joints and global motion representations. Moreover, they typically incorporate partner information solely through cross-attention module, failing to account for the varying influence of the partner across different interaction types. To address these limitations, we propose HHI-Dyn, a novel latent diffusion framework for human-human motion generation. HHI-Dyn first introduces a tailored VAE that effectively fuses global spatial features with local spatial details. At each time step, every local body part adaptively integrates with the global spatial context, yielding a more stable latent representation for subsequent diffusion. A latent diffusion model is then trained to capture this spatial structure conditioned on interaction text. Crucially, we design an interactive expert module for motion pattern learning, where each expert is responsible for capturing a specific type of motion pattern. These experts dynamically adapt to diverse input conditions and partner behaviors, effectively improving both the interactive quality and perceptual realism of the generated two-person interactions. Extensive experiments demonstrate that HHI-Dyn achieves state-of-the-art performance on InterHuman and InterX datasets, with significant improvements in text-motion alignment and motion realism. We will release code and models to facilitate reproducibility.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.