acceptodds
Under review as a conference paper at ICLR 2027

LLMDT: LLM-Based Dual-Transformer for Scene-Aware Text-Driven 3D Human Motion Generation

Abstract

Text-driven 3D human motion generation has a wide range of applications across various domains, while scene-aware generation requires joint reasoning about language, objects, navigation, and human-scene interactions. However, existing methods often struggle to separate global scene reasoning from local motion generation, resulting in inaccurate object grounding, trajectories, and human-scene interactions. We propose LLMDT, a hierarchical framework that addresses spatial reasoning, scene consistency, and motion quality in scene-aware human motion generation. We introduce LLM-based spatial reasoning and object-centric scene understanding to establish structured relationships between human actions and scene elements to improve spatial accuracy. We develop a trajectory-aware generation strategy that integrates global navigation with local scene context to improve scene consistency and generate feasible motion trajectories in complex environments. We further propose a pose-aware motion generation strategy with hierarchical residual representations to improve motion quality by generating realistic and temporally coherent human motions. Experiments on HUMANISE and HumanML3D demonstrate improvements over existing methods in spatial accuracy, scene consistency, and motion quality.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.