acceptodds
Under review as a conference paper at ICLR 2027

Not All Bands Need All Modalities: Frequency-Gated Multimodal Fusion for Time Series Forecasting

Abstract

Multimodal time series forecasters augment numerical inputs with text and image views, but existing designs absorb both auxiliary views indiscriminately: regardless of what these views actually describe, one fusion decision is applied to the whole series and therefore to every frequency component alike. We argue that fusion should instead be selective in the spectral domain. For example, the trend and seasonality summarised by text describe slow, low-frequency structure. High-frequency components should therefore absorb less from the text view than low-frequency ones, since fusing it across the whole spectrum can inject noise into the forecast. To address this issue, we propose **FreqGate**, a frequency-gated multimodal forecaster that assigns fusion weights to each frequency band of the input series. Soft Frequency Partitioning first learns a soft assignment of the spectrum of each input series to low-, mid- and high-frequency bands, rather than imposing a fixed partition shared by all series. Band-Gated Fusion then maps band energies to gates for the text and image encoders. Within each band, numerical tokens attend to the encoders through cross-attention weighted by the corresponding gates. As a result, each auxiliary view contributes only to the bands it can express, learned per series from its spectrum alone. Extensive experiments on eight benchmarks under full-shot, few-shot and zero-shot regimes demonstrate that FreqGate outperforms fifteen existing forecasters.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.