acceptodds
Under review as a conference paper at ICLR 2027

On the Divergence of Accuracy and Mechanism Consistency in Time Series World Models

Abstract

A time series world model (TSWM) takes the observed history of a controlled system and a plan of future actions and exogenous inputs, and predicts how the state will respond. Current approaches build such models as forecasters with the actions as covariates, trained and evaluated on prediction error under the executed plan. Yet a world model is used to compare plans that were never executed, so its response to a changed plan matters as much as its error, and this practice never tests it. Two questions are therefore open: which design choices matter, and whether an accurate forecaster responds to a changed plan the way the real system does. We address both with a formalization of TSWMs and a benchmark. The formalization separates state, actions and exogenous inputs, distinguishes continuous, mode and event actions, and introduces mechanism consistency, a metric built on declared mechanisms, relations between an action and a state variable whose direction is known, such as a vasopressor raising blood pressure: it checks whether shifting an action moves the forecast in the declared direction. The benchmark consolidates eight public datasets with real actions from engineered infrastructure and clinical care, and varies the prediction space, the plan fusion and the plan encoding across seven backbones and five seeds. First, a frozen latent prediction space lowers MAE by 9.9% over observation space and gated fusion at the output by 12.7% over concatenation at the input, both on average and on all eight datasets, whereas encoding the plan over time changes the average MAE by at most 2.2%. Second, prediction error and mechanism consistency diverge: the configuration with the lowest error is at or below chance in consistency on four of the five datasets with declared mechanisms, and no design choice avoids this. Finally, directional supervision, a loss term that penalizes the wrong-signed part of the response to a shifted action, raises consistency significantly on penalized mechanisms with no change in MAE. Together they give TSWMs a recipe: a frozen latent space and output-side fusion for accuracy, and a training objective for mechanism consistency.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.