acceptodds
Under review as a conference paper at ICLR 2027

When and Why Adversarial Training Improves Time-Series Forecasting

Abstract

While adversarial training exhibits contrasting clean-accuracy effects across vision and language, whether and why it benefits unperturbed data in time-series forecasting remains an open question. In this work, we uncover an unexpected within-domain clean-effect reversal: standard raw-history adversarial training (AT) reduces unperturbed test MSE by up to for expressive nonlinear forecasters, yet systematically degrades compact linear models. Through exact finite-network risk identities and asymptotic expansions, we trace this divergence to an intrinsic competition between prediction-variance contraction and parameter displacement: while loss-guided updates act as a frequency-selective regularizer that dampens low-frequency prediction variance, fitting perturbed histories pulls parameters away from the clean empirical minimum, incurring a displacement penalty that can outweigh variance gains in low-capacity architectures. To govern this trade-off, we propose Displacement-Guided Adversarial Training (DGA). DGA introduces an elastic parameter anchor to a pre-trained empirical risk minimization (ERM) checkpoint, effectively suppressing weakly constrained parameter drift while preserving beneficial variance regularization. Across 56 forecasting configurations spanning eight architectures and seven benchmarks, validation-selected DGA outperforms paired ERM in 51 settings (), turning linear-model degradations into consistent clean gains (e.g., MSE on DLinear). Furthermore, DGA transfers effectively to language fine-tuning and vision classification, establishing parameter displacement control as a general stabilizer for adversarial optimization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.