FACETS: Selective Cross-Band Supervision for Robust Time-Series Representation Learning
Abstract
Noise in time-series self-supervised learning can corrupt not only observations but also the supervisory relation used to learn from them. When nuisance is predictable, reconstructable, or preserved across views, standard pretext objectives may consistently favor such nuisance patterns, causing them to become readily decodable from the learned representation. We call this failure mode supervisory contamination. FACETS treats robustness as a problem of supervision design: it makes selective agreement among frequency components of the same sequence the primary learning signal, down-weights weakly supported components via detached relative-regularity estimates, discounts batch-shared drift through directional rectification, and uses temporal prediction only as an auxiliary local constraint. We characterize when component agreement can improve the deployed original-view encoder by separating observed cross-band dependence from nuisance-mediated dependence and the component-to-deployment gap. A controlled diagnostic provides mechanistic evidence: prediction-only supervision makes injected interference highly decodable, whereas cross-band-only supervision sharply reduces nuisance decodability, at some cost in semantic accuracy. Across five biomedical and industrial datasets, FACETS improves 10% label linear-evaluation accuracy by 2.07–4.44 percentage points over the strongest dataset-specific SSL baselines, achieves the best weighted F1 on all five datasets, leads all six ECG transfer directions, and remains strongest under increasing Gaussian corruption. Together, these results show that robust self-supervised learning depends not merely on producing cleaner views, but on controlling which relations are trusted to organize the representation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.