LLMDT: LLM-Based Dual-Transformer for Scene-Aware Text-Driven 3D Human Motion Generation
Abstract
Text-driven 3D human motion generation has a wide range of applications across various domains, while scene-aware generation requires joint reasoning about language, objects, navigation, and human-scene interactions. However, existing methods often struggle to separate global scene reasoning from local motion generation, resulting in inaccurate object grounding, trajectories, and human-scene interactions. We propose LLMDT, a hierarchical framework that addresses spatial reasoning, scene consistency, and motion quality in scene-aware human motion generation. We introduce LLM-based spatial reasoning and object-centric scene understanding to establish structured relationships between human actions and scene elements to improve spatial accuracy. We develop a trajectory-aware generation strategy that integrates global navigation with local scene context to improve scene consistency and generate feasible motion trajectories in complex environments. We further propose a pose-aware motion generation strategy with hierarchical residual representations to improve motion quality by generating realistic and temporally coherent human motions. Experiments on HUMANISE and HumanML3D demonstrate improvements over existing methods in spatial accuracy, scene consistency, and motion quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.