acceptodds
Under review as a conference paper at ICLR 2027

Multimodal Clustering for Weak Annotation of Human Activity Recognition

Abstract

Annotating Inertial Measurement Unit (IMU) data for Human Activity Recognition (HAR) is challenging and time-consuming, as raw inertial signals are difficult for humans to directly interpret and assign activity labels. Recent studies therefore leverage synchronized ego-centric video for weak annotation, using visual representations to cluster activity samples and requiring manual labels only for representative samples. However, visual data can raise privacy concerns and its representations may encode scene and recording context unrelated to the performed activity, influencing the resulting feature space, clustering structure, and activity labels. To incorporate additional activity-relevant information, we propose WAFER (Wearable And visual Foundation Embeddings for weak annotation of HAR), a multimodal weak annotation method that incorporates representations from a pretrained wearable foundation model alongside visual and optical-flow features. Temporally aligned feature embeddings from video, wearable IMU signals, and optical flow are independently standardised, weighted, and fused before clustering. Human annotation is required only for few representative samples, after which activity labels are propagated to the remaining samples within each cluster. The resulting weak annotations are then used to train downstream HAR models using only IMU data, such that video is not required during downstream training. Across three benchmark datasets, WAFER improves weak annotation accuracy by up to 6.7. Using these weak annotations for downstream HAR training improves accuracy by up to 6.7 and macro F1 by up to 10.5, achieving performance approaching the performance of fully supervised training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.