acceptodds
Under review as a conference paper at ICLR 2027

Bridging Semantic Reasoning and Geometric Refinement with Diffusion Language Models for Multimodal 3D Human Pose

Abstract

Multimodal 3D human pose modeling must jointly interpret images, language, and poses while producing anatomically coherent 3D configurations. Existing autoregressive models struggle to preserve coupled joint constraints under causal decoding, whereas continuous pose diffusion priors improve geometric plausibility but lack language grounding. Masked discrete diffusion language models (dLLM) offer a promising alternative through bidirectional context and iterative revision, which are naturally suited to infilling and constrained modification. Motivated by this property, we propose PoseDLLM, a unified discrete-to-continuous diffusion framework that decomposes pose prediction into two complementary stages. DLLM serves as the semantic engine, leveraging bidirectional attention and holistic parallel denoising to handle multimodal conditions across all task configurations. A continuous diffusion prior then refines the coarse prediction toward the biomechanically plausible pose manifold via test-time optimization. Extensive experiments on Human3.6M, PoseScript, PoseFix, and Posepart benchmarks demonstrate that PoseDLLM consistently outperforms autoregressive LLM baselines across all task configurations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.