C4TS: Co-Encoding and Co-Decoding via Temporal-Visual Cross-Modal Collaboration for Time Series Forecasting
Abstract
Visual representations of time series have been widely explored to improve forecasting. In temporal–visual feature-fusion approaches, interaction between parallel temporal and visual encoders feeds a single temporal decoding pathway, yet does not inherently reconcile the structural differences between temporal token sequences and spatially organized visual features. In temporal–visual prediction-fusion approaches, parallel decoding preserves both forecasting pathways, but globally shared weights cannot adapt their combination to individual input instances and future positions. We propose Co-Encoding and Co-Decoding via Temporal–Visual Cross-Modal Collaboration for Time Series Forecasting (C4TS), a framework that coordinates two forecasting pathways through both collaborative encoding (co-encoding) and collaborative decoding (co-decoding). For co-encoding, Spatial-Aware Visual Adaptation bridges the structural gap between temporal and visual representations by adapting spatially organized visual features into a cross-attention memory, enabling visual features to enhance temporal representations. For co-decoding, Instance × Future-Patch Adaptive Fusion assigns distinct fusion weights to each input instance and future patch, adaptively combining temporal predictions with visual reconstruction forecasts. These weights are generated by learnable future-patch queries that retrieve historical evidence separately from the temporal and visual representations. Experiments on public time series benchmarks demonstrate that C4TS achieves state-of-the-art overall forecasting performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.