Beyond Token Alignment: Human-Centered Posterior Learning for Multimodal Human Sensing
Abstract
Multimodal human sensing combines heterogeneous sensors such as RGB, depth, LiDAR, mmWave radar, WiFi CSI, and RFID for robust and privacy-preserving perception. Existing methods typically align or aggregate modality tokens in a shared embedding space. However, these tokens arise from distinct physical mechanisms and are differently entangled with environment factors, sensing noise, and uncertainty. We therefore argue that the common object across modalities is the underlying human state rather than token correspondence. Based on this insight, we propose HUPER, a Human-centered uncertainty-aware Posterior Estimation and Refinement framework. HUPER preserves sensor-specific inductive biases with modality-aware encoders, estimates a probabilistic human posterior for each modality, separates human content from environment factors, and forms a reliability-calibrated posterior consensus. It further corrects residual uncertainty through an uncertainty-conditioned vector field in the human state space. Experiments on MM-Fi and XRF55 demonstrate strong performance across pose estimation, activity recognition, missing modalities, and heterogeneous RF settings. On MM-Fi HPE, HUPER-Flow reduces full-modality and partial-average MPJPE over X-Fi by 22.7% and 15.9%, respectively. It improves full-modality MM-Fi HAR accuracy by 13.3 percentage points. On XRF55 HAR, it achieves 97.1% full-modality accuracy and improves the full- and dual-modality results over X-Fi by 7.3 and 13.7 percentage points. These results establish human-centered posterior learning as a more reliable alternative to direct token alignment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.