acceptodds
Under review as a conference paper at ICLR 2027

Think with Motion: A Unified Understanding and Generation Framework for Individual and Interactive Human Motion

Abstract

Text-driven human motion modeling bridges high-level semantics with complex physical dynamics. For interactive motion generation, prevailing methods leverage coupled representations in the raw motion space, struggling to interpret underlying social intent through interactive text-conditioned descriptions. Meanwhile, emerging motion–language models unify motion understanding and generation, whereas the shared parameters across both tasks induce severe interference. To address these issues, we propose , a simple and unified architecture for individual and interactive motion understanding and generation. Specifically, we introduce a relation-decoupled representation that factorizes interactions into individual movements and relative geometry. further introduces a task-specific Mixture-of-Transformers architecture to mitigate interference between motion understanding and generation. We additionally establish a participant-level Chain-of-Thought reasoning mechanism, enabling motion understanding to explicitly guide interactive motion generation. Extensive experiments on HumanML3D and InterHuman datasets demonstrate that achieves superior performance across all motion modeling tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.