SFCATime: Soft Spectral Representation and Adaptive Cross-Modal Fusion for Time Series Forecasting
Abstract
Multivariate time-series forecasting can benefit from complementary frequency-domain and language-derived representations, yet their effectiveness depends on how frequency structure is organized and how semantic information interacts with numerical features. A shared spectral projection leaves frequency specialization to downstream encoders, while standard cross-modal attention relies primarily on query–key affinity to determine numerical–semantic interactions. We propose SFCATime, an adaptive spectral–semantic forecasting framework that addresses both issues. First, SFCATime learns soft, overlapping assignments of frequency bins to latent spectral components and uses them to construct frequency-dependent corrections to a retained full-spectrum pathway. This preserves the complete magnitude spectrum while introducing spectral specialization without hard frequency boundaries. Second,AdaptiveMask Cross-Modal Attention augments base numerical–semantic affinity with a learnable signed relation correction before attention normalization, while adaptive branch aggregation and residual modulation separately control the contribution of the aligned representation. The pretrained language model remains frozen throughout forecasting training. Across eight widely used multivariate forecasting benchmarks, SFCATime achieves the best or tied-best average MAE on all eight datasets and the best average MSE on six. Compared with the most directly related multimodal baseline, it obtains lower average MSE and MAE on seven datasets each. Ablation studies and representation analyses further support the complementary effects of adaptive spectral representation and relation-aware cross-modal regulation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.