Which Channels, or How Many? Matched Controls for Data-Derived Attention Masks in Multivariate Forecasting
Abstract
Forecasters that mix channels with attention are often given a fixed relational mask estimated from data, and the benefit is usually reported against the unmasked model. Such comparisons change more than the graph: implementations, training runs and machines differ, and a mask also makes attention sparse, so a gain does not show that the estimated edges matter. We study this question for three masks estimated from road-sensor series (partial correlation, Pearson correlation and lead–lag communities, assigned to separate heads of an inverted Transformer), using two controls trained in the same launch as the masked model: an all-ones mask that leaves every parameter and setting unchanged, and a rewiring of each relation that keeps every sensor's degree. We find that the masks lower test error against the all-ones control in all 22 dataset–horizon cells on six road networks, by 13% on average, and that the gain grows with the forecast horizon on every network; on three non-road series they do not help. We find that degree-preserving rewiring keeps only the smaller part of the gain, so most of it requires the estimated edges rather than sparsity alone. We also find that pairing runs from different launches moves the estimated effect by several points, more than the spread between seeds, which is why the controls must be trained together. We release the masks of the rewiring study, estimated without the test split, with their builder, a per-run configuration table that names each run's mask file, and the analysis code.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.