acceptodds
Under review as a conference paper at ICLR 2027

Refine Together: Mutual Pose-Kinematic Feedback for 3D Human Pose Estimation

Abstract

Mamba has shown promise for modeling complex spatiotemporal relations in 3D human pose estimation (3D-HPE). However, ambiguous 2D observations can correspond to different 3D configurations, while structural and motion relations remain implicit in latent pose representations. Explicit 3D kinematic relations provide direct target-space cues, yet their reliability depends on the accuracy of the estimated pose. To address these limitations, we propose CoKinPose, a target-space 3D pose modeling framework that jointly refines pose and kinematic representations within Mamba. First, a kinematics-driven state update mechanism derives skeletal and motion relations from an initial 3D pose estimate and incorporates them into feature processing and selective state evolution. A confidence-guided gate adaptively regulates their contributions using 2D observation confidence and the current kinematic representations. Second, an iterative pose-kinematic mutual refinement mechanism applies this state update across successive Mamba blocks and feeds updated pose representations back to refine the kinematic representations used by subsequent blocks. Experiments on Human3.6M and MPI-INF-3DHP demonstrate consistent improvements under standard and challenging conditions, with further analyses validating enhanced structural and motion modeling under spatial skeletal and temporal motion ambiguities. Our code and models will be made publicly available upon publication.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.