Joint Visual-Temporal Diffusion Model for Probabilistic Time Series Imputation
Abstract
Score-based diffusion models have emerged as powerful probabilistic tools for time series imputation, but their denoising process is still primarily conditioned on observed temporal values. When missing values are dense or form long contiguous gaps, such temporal context becomes sparse and fragmented, weakening the structural guidance available for generation. We revisit this problem from a complementary perspective: even when numerical observations are incomplete, the macroscopic shape of a time series often remains visually discernible in its line-graph rendering. Rather than treating imputation as an image-generation problem, we use this visual representation as a structural prior for temporal diffusion. We propose VT-Diff, a Joint Visual-Temporal Diffusion framework for probabilistic time series imputation. VT-Diff renders multivariate time series as variable-aware line graphs and encodes them with a Vision Transformer to obtain global and local structural tokens. The central component is a Cross-Attention Bottleneck with learnable temporal anchors, which translates spatial visual tokens into time-step-wise conditions for the temporal denoising network. This design allows visual cues to guide the global shape and local variations of the imputed trajectory, while the diffusion model performs generation in the original numerical time-series domain. Extensive experiments on diverse missing patterns show that VT-Diff achieves competitive or superior performance against strong baselines and produces more structurally coherent imputations, especially under extreme missingness.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.