acceptodds
Under review as a conference paper at ICLR 2027

Multimodal Foundation Priors for Diffusion-Calibrated Zero-Shot Time Series Reconstruction

Abstract

Zero-shot time series reconstruction asks a model trained where a target is observed to reconstruct it where it never was, from exogenous inputs alone. Prior work builds an informed prior from a single exogenous source and calibrates it with diffusion, so the prior is only as informative as that source. We ask whether a multimodal foundation model can supply a richer prior. We pretrain a transformer by masked reconstruction over heterogeneous inputs, daily gridded fields, station-scale series, and static attributes, tokenized at their native resolutions and fused by joint space-time attention, with a query decoder that returns a representation at any coordinate, day, variable, and depth, including locations that have no labels. Downstream, this representation enters a recurrent prior model through a small learned map, and a conditional diffusion calibrator refines the prior into a predictive distribution; prior, map, and calibrator are fine-tuned jointly. On four zero-shot reconstruction benchmarks (streamflow, stream temperature, soil moisture, and dissolved oxygen, under spatial cross-validation), the multimodal prior improves two tasks, with the largest gains where the single-source prior is weakest, and we characterize why the other two do not benefit: their target statistics at unseen locations are not identifiable from any input. Ablations show that the width of the representation map matters little, that fine-tuning it jointly with the diffusion calibrator matters a great deal, and that the point read-out of the calibrator must be chosen with care.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.