ISIMUD: A Unified Diffusion Model for Time Series Captioning and Generation
Abstract
Time series rarely appear in isolation in real-world, they are often accompanied by texts that provide context and expert knowledge. This naturally gives rise to two complementary cross-modality tasks: time series captioning and text-conditioned generation. Although these two tasks correspond to opposite conditional directions of the same underlying text-time series joint distribution, existing methods typically model them separately, limiting their ability to exploit shared cross-modal representations. Moreover, prior approaches are often limited to specific datasets or domains, due to the scarcity of high-quality, semantically aligned text–time series pairs. To address these challenges, we present ISIMUD, a multimodal foundation model that unifies both tasks within a shared diffusion-based architecture, enabling native modeling of both time series and text modalities. In addition, we introduce BOOK-OF-ISIMUD, a large-scale cross-domain dataset of text-time series pairs with domain-agnostic semantics, providing broad distributional coverage for training and evaluating multimodal time series models. Extensive experiments demonstrate that ISIMUD performs effectively in both in-domain and zero-shot settings, while adapting efficiently to multiple downstream tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.