CDPose: Compositional Diffusion for Whole-Body Pose
Abstract
Generating 3D whole-body poses that simultaneously capture body articulation, detailed hand configurations, and expressive facial motion remains a challenging task due to joint rotations evolving in an unconstrained space, where small local errors propagate and escalate across parts. To address these limitations, we propose CDPose, a diffusion-guided whole-body pose model that represents each articulated joint in 6D rotation and converts it inside the generator to an axis–angle before kinematic evaluation. The denoiser is trained under a sub-variance-preserving stochastic differential process with a signal-to-noise–weighted objective. During inference, CDPose produces a single-step Tweedie estimate of the clean pose, with an optional short reverse-time refinement for additional sharpening. The whole-body assembly combines part-specific diffusion heads with a fused residual head that learns cross-part coordination, such as mirrored hand motion and face–torso co-articulation. The coverage and diversity are further improved by mixed training on heterogeneous whole-body and part-only data, using selective masking and supervision. Comprehensive experiments demonstrate that CDPose is robust and versatile across diverse benchmarks spanning body, hand, face, and whole-body pose modeling. In particular, CDPose achieves the lowest errors on the ARCTIC dataset, obtaining a PA-MPJPE of 26.92 mm for body mesh recovery, a PA-MPVPE of 7.10 mm for hands, 2.51 mm for the face, and 23.24 mm for whole-body mesh recovery (PA-MPVPE over all vertices). The code will be publicly released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.