Temporal Operator Attention: Recovering Common Operators in Time-Series Analysis
Abstract
Real-world time-series applications repeatedly exhibit five common structural regimes: seasonal extrapolation, trend/anomaly separation, cross-channel demixing, regime-shift alignment, and transient recovery. Each requires a signed or zero-sum linear map over time, not a convex combination of the sequence's own values, and some further require this map to adapt to the input. Linear and MLP-based mixers, despite far fewer parameters, often match or outperform attention-based forecasters on exactly this structure, because they mix directly along the temporal dimension via a signed, position-indexed operator, rather than one derived from content-based similarity and confined to a probability simplex. We prove that softmax attention's row-stochastic kernel where nonnegative weights summing to one cannot realize any of the five recurring regimes from a single head, tracing the gap to this primitive-level mismatch rather than model scale. We propose Temporal Operator Attention (TOA), which restores this signed, position-indexed operator to the attention pathway itself: learnable dense residual operators wrap the score matrix, supplying signed degrees of freedom, while a gated variant further makes the operator input-adaptive. Stochastic Operator Regularization (SOR) controls overfitting from this added capacity by regularizing only the learned offsets, preserving identity-based signal transport. Across Autoformer, PatchTST, iTransformer, and DUET, TOA consistently improves long- and short-term forecasting, anomaly detection, and classification, with the largest gains where signed temporal mixing matters most.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.