acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Multimodal Learning for Large Language Models in Time Series Forecasting: The Problem of Pseudo-Multimodality

Abstract

LLM-based time-series forecasting increasingly couples numerical predictors with prompts generated from the same observed series. Although the resulting branches differ in form, neither receives independent sample-specific evidence. They are two representations of a common observational source, a condition we call pseudo-multimodality. Cross-view similarity alone does not establish that the language representation should alter a numerical forecast. We formulate sequence-derived LLM forecasting as a conditional correction problem: given a complete numerical prediction, should a prompt-derived candidate modify it for each sample and variable? Conditional Incremental and Robust Fusion (CIRF) uses the numerical forecast as the default prediction and the language branch as a rejectable correction path. It removes the prompt component predictable from the numerical representation, decodes the remaining conditional residual as a candidate correction, and applies it through a utility-supervised gate. Training also limits the excess perturbation sensitivity introduced by this optional path when both representations inherit the same corrupted observation. We evaluate CIRF under matched input and forecasting-horizon protocols on nine public benchmarks. Across six same-checkpoint routing settings, selective correction yields lower reported MSE and MAE than mandatory correction, while language-route activation ranges from 0.2% to 77.5%.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.