Learning Representations for Multivariate Time Series Classification via Context Prediction
Abstract
Self-supervised representation learning for multivariate time series classification is dominated by contrastive learning and masked modeling. However, these paradigms incur substantial computational overhead and often rely on complex data augmentations or reconstruction decoders that introduce domain-specific biases. In this work, we propose **C-Pred**, a lightweight self-supervised framework based on next-patch context prediction. C-Pred formulates self-supervised learning as a -way classification problem: given a reference patch, the model identifies its immediate temporal neighbor from candidate patches within the same instance. This lightweight formulation, requiring only patch-level encoding rather than full-sequence processing, makes **channel-independent encoding** practical: each channel is processed independently by a shared univariate backbone, allowing a single model to accommodate arbitrary channel dimensions without increasing the backbone size. By requiring only a small number of temporal patches and no data augmentation, cross-instance negative sampling, or signal reconstruction, C-Pred substantially reduces the computational cost of self-supervision. Across six large-scale benchmarks (20k instances) and 28 UEA multivariate datasets, C-Pred matches or exceeds state-of-the-art representation quality while achieving up to a **7.5 speedup** in training time over competitive baselines. On the large-scale benchmarks, channel-independent encoding further improves accuracy by 2.99 percentage points over channel mixing, achieving 84.36% average accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.