acceptodds
Under review as a conference paper at ICLR 2027

WhatIf-Bench: Benchmarking Intervention-Aware Time Series Forecasting

Abstract

Recent time series foundation models (TSFMs) can consume covariates in-context, extending zero-shot forecasting to problems driven by external factors. This enables scenario analysis: forecasting how a target responds to events that set its covariates to chosen future values. Whether they can do this reliably is an open question. We introduce WhatIf-Bench, a benchmark measuring how well a forecaster uses in-context covariates to forecast a target's response to future events. WhatIf-Bench is semi-synthetic: realistic time series come from mechanistic simulators of data centers, power grids, road traffic, and water networks, on which we impose controlled covariate-target relationships grounded in each system. Each task presents a history of events that intervene on covariates, then asks to forecast a future event that either replays seen covariate values (a Replay future) or sets them to unseen ones (a Novel future). Varying the number of covariates, historical events, and event length yields 37,665 tasks. Evaluating covariate-aware TSFMs and frontier LLMs, we find the LLMs reach the highest skill scores (relative improvement over a naive baseline), ahead of the best TSFMs. Across the axes, however, every model forecasts well on Replay futures (65–85% skill) but drops to 30–50% on Novel ones. Skill also falls as events set more covariates jointly. More historical events raise skill across the board, especially on Novel futures, but large differences between models remain. WhatIf-Bench offers a testbed toward models that support arbitrary scenario analysis.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.