acceptodds
Under review as a conference paper at ICLR 2027

Text-Conditioned Time-Series Generation: Cross-Modal Alignment and Fine-Grained Control in Time-Frequency Image Space

Abstract

Text-conditioned time-series generation enables flexible and fine-grained control over generated time series, supporting a wide range of applications. Despite its potential, existing methods face two key challenges: (1) limited alignment between time series and text (TS-text alignment); and (2) limited adherence to fine-grained textual instructions. In this paper, we propose TextPlot, an image-based diffusion framework for text-conditioned time-series generation. To address the first challenge, we introduce the short-time Fourier transform (STFT) as an intermediate representation, mapping a one-dimensional time series to a two-dimensional time-frequency image with explicit temporal and frequency coordinates. The resulting time-frequency image provides a structured representation for relating textual descriptions to temporal dynamics. For the second challenge, TextPlot employs a dual-path text-conditioning mechanism. Global pooling provides sentence-level context, while token-level cross-attention preserves fine-grained textual information. These two conditioning paths jointly guide a 2D U-Net to generate time-frequency images through iterative denoising. Across one synthetic and four real-world datasets, TextPlot achieves superior TS-text alignment compared with the evaluated baselines. Further experiments show average relative improvements of 79.1% and 83.5% over the evaluated baselines in following original and edited textual requirements, respectively.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.