Separating Command Error, Chaining Cost and One-Call Error in the Rollout Evaluation of a Video World Model
Abstract
Action-conditioned video world models are usually evaluated by rolling them out under logged actions and plotting the error against horizon, and the growth of that error is attributed to errors compounding through the model's own outputs. The rollout score alone cannot support that attribution, because it has no reference: it does not show whether the evaluation gave the model the right commands, how much of the growth a single call to the same moment would already show, or whether the actions mattered. For models that take a time shift as input, three additional generations on the same items provide these references without retraining: a rollout with the commands recomputed as in training, one call to each scored horizon, and a rollout that repeats the first command. We applied them to the public evaluation of a released family of navigation world models (NWM; 49M to 1.01B parameters), on fresh test cohorts and with decision rules registered in advance. Correcting the released rollout code, which gives every call after the first a wrong command, lowers the 1B model's 4 FPS rollout LPIPS by 18.8% [16.9, 20.6]; in a post-hoc analysis the growth of the command error accounts for a third of the released curve's rise from 2 to 8 s. The release's authors attribute a crossover between their 1 FPS and 4 FPS rollouts to accumulated errors and loss of context. On the release's own 20-trajectory protocol, the released 1B checkpoint shows this crossover under the released commands only in the point estimates, and it disappears once the commands are corrected (post hoc; we cannot tell whether their figure used this code). With correct commands, the growth of the error that chaining adds over a single call accounts for between a sixth and two fifths of the rise from 2 to 8 s, and the two cohorts' estimates of that growth disagree beyond their intervals (post hoc). The larger share's post-hoc interval reaches about one half, so removing this chaining cost while leaving the one-call error unchanged would leave at least about half of the rise. Scale lowers the one-call error, although the released 1B checkpoint is worse than the 700M one on single calls (post hoc). The effect of scale on the chaining cost is unresolved on LPIPS; on DreamSim the cost decreases. We recommend reporting rollout scores together with these references.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.