ZEST: Zero-Shot End-to-End Time Series State Detection
Abstract
Time series state detection (TSSD) identifies state boundaries and recurring states in multivariate sequences. Deep self-supervised TSSD methods typically combine target-specific representation learning with clustering, hindering direct transfer to unseen sequences. Our analyses identify persistent conflicts in temporal-proxy supervision and limited transfer gains from enlarging the pretraining corpus under the tested fixed-budget settings. We introduce ZEST, a zero-shot, end-to-end framework that recasts TSSD as hierarchical structured generation, adapting a pretrained vision-language model to predict state structure from global sequence images. To build scalable supervision without manual annotation, we use forecast consistency to guide state-prototype construction from large-scale unlabeled real-world time series. A controllable scheduler then recombines these prototypes into multivariate sequences with exact state labels. To capture state relations across granularities, we construct hierarchical supervision for joint generation of leaf segmentations and state-merge trees. The adapted model provides a default leaf partition and multiple coarser candidates on unseen sequences without target labels, parameter updates, hyperparameter tuning, or clustering. Experiments on six benchmarks demonstrate strong zero-shot leaf performance, while a separate targetlabel-assisted study shows the utility of selecting among the predicted granularities.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.