MUST: Multimodal Unified Sensor Tokenization
Abstract
The development of general time-series foundation models is fundamentally constrained by the asynchronous, multi-rate nature of physiological data. Traditional early-fusion architectures attempt multimodal integration by forcing diverse streams onto rigid, synchronous interpolation grids. However, this artificial alignment incurs severe computational overhead, wastes parameter capacity on imputed noise, and causes cross-modal contamination where high-variance signals dominate. To overcome these limitations, we introduce Modality-Specific Early Tokenization coupled with semantic latent predictive alignment. Each sensor stream is processed by an independent patch encoder tailored to its native frequency, structurally isolating noise and bypassing artificial data imputation. By dynamically concatenating only valid active tokens, a shared transformer organically aligns multi-rate sequences with routing compute overhead. Evaluated on large-scale, real-world datasets comprising over 3.7 billion minutes of continuous sensor data from 50,000 users across nine modalities and 15 distinct activity classes, our framework consistently outperforms generic tokenization baselines across micro- and macro-level tasks. For fine-grained kinematics, it achieves a relative 18.2% Macro F1 score improvement on complex, long-tail activities. Furthermore, within the Digital Well Being (DWB) dataset, our architecture drives significant improvements in predicting socioeconomic factors related to health (69.7% vs. 65.7% average accuracy) while also achieving a higher average continuous tracking score across five physiological metrics (Pearson vs. ). Ultimately, we establish native modality tokenization as a highly resilient, architecture-agnostic primitive for multi-resolution time-series. To facilitate future research, we will open-source the DWB benchmark.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.