Motion as Common Ground: Aligning Heterogeneous Sensing Modalities for Human Activity Recognition
Abstract
Human activity recognition (HAR) using heterogeneous sensors is challenging because modalities such as RGB, depth, LiDAR, mmWave, and WiFi capture the same activity through different physical signals and produce data with distinct structures. Existing methods typically bridge this heterogeneity through modality-invariant learning or LLM-based semantic alignment. We take a different view: *despite different sensing mechanisms, these modalities observe the same underlying human motion*. Since motion is shared across sensors and directly tied to the activity, it provides a natural basis for cross-modal alignment. In this paper, we introduce **UniMotion**, a motion-grounded framework that first establishes a shared temporal-joint motion space using a skeleton-pretrained motion model. UniMotion is then optimized in two stages. In Stage 1, modality-specific Motion Adapters with learnable joint queries map each sensor's native features into joint-indexed tokens in this space, aligning heterogeneous modalities within a common temporal-joint representation. Skeleton sequences are used only during training. In Stage 2, **Reliability- and Disagreement-Aware Motion Fusion (RDMF)** integrates the aligned modalities for multimodal recognition. RDMF estimates time-varying modality reliability to form a weighted motion consensus and uses cross-modal disagreement to retain complementary evidence beyond the consensus. Evaluated on MM-Fi and XRF55 across seven sensing modalities, UniMotion achieves state-of-the-art average accuracy under random-split, cross-subject, and cross-environment evaluation protocols and remains effective across varying modality combinations. These results support the use of structured human motion as a physically grounded, task-relevant shared representation for multimodal HAR.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.