Vision Pretraining for Time-Series Forecasting Depends on More Than Undirected Grid Topology
Abstract
Vision models can forecast time series by treating numerical sequences as masked images, but why natural-image pretraining transfers remains unclear. We test whether undirected grid structure explains this benefit. Phase, cycle, and joint reflections preserve observed values and undirected adjacency exactly while changing how temporal structure is embedded in the visual coordinate system. Across 11 datasets, 71 series, 2,272 forecast origins, and three masked-image checkpoints, reflections consistently reduce pretrained-weight advantage. The effect remains under architecture-matched random controls and alternative metrics including MAE, MASE, and sMAPE. A post-hoc loss decomposition shows that reflections primarily increase pretrained-model losses while leaving matched random-model losses nearly unchanged. These results show that undirected topology alone cannot explain cross-modal transfer from vision pretraining to time-series forecasting; the roles of direction, absolute position, and boundary anchoring remain unresolved.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.