Fast Rates for Nonstationary Weighted Risk Minimization
Abstract
Weighted empirical risk minimization is a common approach to prediction under distribution drift. This article studies its out-of-sample prediction error under nonstationarity. We provide a general decomposition of the excess risk into a learning term and an error term associated with distribution drift, and prove oracle inequalities for the learning error under mixing conditions. The learning bound holds uniformly over arbitrary weight classes and accounts for the effective sample size induced by the weight vector, the complexity of the weight and hypothesis classes, and potential data dependence. We also analyze the drift error in the square loss setting and provide explicit bounds for structural-break and smoothly drifting processes. We then combine the learning and drift analysis in (auto-) regression problems with linear models, basis approximations, and neural networks. In all cases, the optimal weights are characterized by a bias–variance trade-off between learning and drift, while minimax-optimal rates (up to logarithmic factors) are recovered when specialized to unweighted and stationary settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.