acceptodds
Under review as a conference paper at ICLR 2027

Variable Input Masked Autoencoders for Wrist-Worn Accelerometry

Abstract

Large-scale repositories of wearable sensor data have enabled the development of foundation models for human activities and behavior. However, existing models are pretrained using fixed input durations and sampling rates, despite substantial heterogeneity in recording conditions, downstream tasks, and deployment constraints. We introduce VarMAE, a variable input masked autoencoder that learns a single representation across sampling rates and temporal contexts. During pretraining we construct two views of the same signal: a long duration high frequency view, and a short duration, lower frequency view, obtained via random cropping and downsampling. We perform masked reconstruction within each view and introduce a cross-resolution masked reconstruction objective that uses visible tokens from the long view to predict masked tokens from the short view. Cropping and downsampling impose complementary information bottlenecks, encouraging the model to learn representations that remain effective across input durations and sampling rate changes. We pretrain VarMAE on the large-scale NHANES dataset, producing a single pretrained encoder matching or outperforming controlled foundation model baselines on eight benchmark datasets. With only 0.7M parameters, VarMAE is parameter efficient and shows that a single foundation model can achieve broad transfer capabilities across sensing configurations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.