Bridging Semantic Reasoning and Geometric Refinement with Diffusion Language Models for Multimodal 3D Human Pose
Abstract
Multimodal 3D human pose modeling must jointly interpret images, language, and poses while producing anatomically coherent 3D configurations. Existing autoregressive models struggle to preserve coupled joint constraints under causal decoding, whereas continuous pose diffusion priors improve geometric plausibility but lack language grounding. Masked discrete diffusion language models (dLLM) offer a promising alternative through bidirectional context and iterative revision, which are naturally suited to infilling and constrained modification. Motivated by this property, we propose PoseDLLM, a unified discrete-to-continuous diffusion framework that decomposes pose prediction into two complementary stages. DLLM serves as the semantic engine, leveraging bidirectional attention and holistic parallel denoising to handle multimodal conditions across all task configurations. A continuous diffusion prior then refines the coarse prediction toward the biomechanically plausible pose manifold via test-time optimization. Extensive experiments on Human3.6M, PoseScript, PoseFix, and Posepart benchmarks demonstrate that PoseDLLM consistently outperforms autoregressive LLM baselines across all task configurations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.