ProxiHG: Latent Proximal-joint Hypergraph Encoding for Sparse-IMU 3D Human Pose Estimation
Abstract
Sparse-IMU-based 3D human pose estimation aims to recover full-body motion from a small number of inertial sensors. Under the standard six-IMU setting, key proximal joints, such as the shoulders and hips, are not directly observed. As a result, similar distal IMU signals may correspond to substantially different full-body poses, leading to inherent pose ambiguity. To address this issue, we propose ProxiHG, an encoder-side structural representation framework that explicitly introduces latent proximal joints before pose decoding. Specifically, ProxiHG employs a two-level hypergraph to first model relations among observable IMU nodes and then capture high-order dependencies among sensors, latent proximal joints, and body parts. The resulting structured representation is temporally fused and progressively decoded into full-body motion. Experiments on DIP-IMU and TotalCapture demonstrate consistent improvements in pose accuracy, spatial reconstruction, motion stability, and cross-dataset generalization. These results show that explicitly modeling unobserved proximal joints provides an effective structural inductive bias for sparse-IMU-based full-body motion estimation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.