Beyond Frame Shuffling: Interpreting Temporal-Order Sensitivity in Video Prediction
Abstract
A video predictor's response to reordered frames is difficult to interpret: reordering can change local transitions and perturbation magnitude, while aggregate error scores can conceal changes in the predictions themselves. We study this ambiguity through controlled comparisons with explicitly preserved input properties. Our construction enumerates six histories with identical frame content, boundary frames, and final five observations, and groups them by their directed adjacent-frame-pair multisets. Within-group pairs preserve local transitions; between-group pairs provide references with the same frame content. We evaluate all pairs using prediction disagreement, pixelwise error change, and aggregate error change. Experiments comprise 896 forecasts on 32 Moving MNIST clips under unmasked and masked observations, using released SimVP and OpenSTL ConvLSTM predictors. For unmasked first-frame predictions, SimVP's transition-matched pairs exhibit 14.7% greater mean prediction disagreement than unmatched pairs. Post-hoc comparison against references with exactly equal input L1 distance reduces this gap to 6.3%, with a similar result under stricter per-pixel distance matching. The ConvLSTM predictor remains close to parity, and some masked comparisons reverse the ordering. Output-range decomposition and stepwise analysis further distinguish prediction changes from those retained by scalar errors. Together, the controls and decompositions turn order sensitivity into a precisely specified comparison: the input constraints define what is held fixed, and the readout defines which prediction changes remain visible.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.