MoJEPA: Learning Transferable Motion Representations via Structured Dual-View Joint-Embedding Prediction
Abstract
A fundamental challenge in sensor-based Human Activity Recognition (HAR) is learning transferable motion representations from unlabeled inertial sensor signals for effective cross-domain adaptation. Existing self-supervised methods often rely on signal reconstruction or consistency across augmented views, which may capture low-level signal variations rather than the underlying physical structure of motion, limiting the transferability of the learned representations. To address this problem, Joint-Embedding Predictive Architectures (JEPAs) offer a framework for learning representations through latent-space prediction. Building on this framework, we propose MoJEPA, a dual-view joint-embedding predictive framework that learns transferable motion representations by explicitly modeling the underlying motion structure. MoJEPA introduces two structured masking views along the channel and temporal dimensions, guiding the model to predict masked latent representations from complementary motion cues. The channel view encourages the representation to capture cross-channel physical dependencies, while the temporal view captures temporal motion dynamics. Jointly optimizing these complementary objectives encourages representations that capture shared motion structure. Extensive experiments with pretraining on two large HAR datasets and evaluation across seven target datasets show that MoJEPA outperforms existing self-supervised baselines under frozen-encoder evaluation with few-shot and low-label adaptation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.