Mobile-Human-Bench: A Unified Benchmark for Generalizable Mobile Human Sensing
Abstract
Mobile sensing enables modeling of human states and behaviors, yet progress is difficult to compare across studies because datasets differ in sensing modalities, targets, temporal resolution, preprocessing, and evaluation. We introduce Mobile-Human-Bench, a unified benchmark for evaluating generalization across time, users, and independently collected studies. It harmonizes seven longitudinal studies spanning 1,751 participants across nine countries and yields 295,162 labeled feature rows after benchmark construction. The benchmark operates at the structured sensing-feature level, supports momentary, daily, and weekly prediction, and covers affective, mental-health, and behavioral targets. We evaluate classical machine learning, deep tabular models, domain generalization, unsupervised domain adaptation, and tabular foundation models under within-user temporal, unseen-user, and cross-dataset evaluation. Results show that model rankings and apparent gains change substantially across generalization settings, targets, and timescales. Tabular foundation models are strong in within-user settings but do not retain a systematic advantage for unseen users or independently collected studies, while standard domain generalization and adaptation provide limited gains over matched controls. Mobile-Human-Bench provides a reproducible testbed for identifying which modeling advances remain robust across users, studies, outcomes, and timescales. Resources are available at https://anonymous.4open.science/r/Mobile-Human-Benchmark-B3C5.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.