The Price of Not Forgetting: A Covering-Based Memory-Error Analysis for Continual Time-Series Forecasting under Regime Drift
Abstract
A forecaster running on a live stream faces an old dilemma: adapting to a new regime can erase what it knew about the old one. The usual escape is to freeze what has been learned and grow a fresh expert per regime - forgetting becomes impossible by construction, but the guarantee is bought with memory, and nothing says how much, or when to stop growing. We ask what can be certified about that bargain. The regimes a forecaster serves form a small metric space of distributions, and it is the intrinsic dimension of that space, not the raw channel count, that governs how error falls as experts are added. Covering arguments turn expert allocation into a budgeting problem with explicit upper certificates: for the regimes observed, and, under stated assumptions, across the whole condition space. On one narrow, fully specified class of drifts we also match the rate from below over a limited range of budgets - a conditional match, not universal optimality, and our slack terms are upper bounds rather than proved error floors. Three design rules follow: how large an adapter must be, when adding experts stops paying, and which regimes to keep in one pass. On five real datasets, freezing gives exactly zero measured forgetting - because old-regime parameters never move - yet does not uniformly out-learn methods that tolerate some. Low-rank adapters help only behind a bottleneck; a cross-channel drift statistic repairs a real blind spot in a controlled test but not on real streams; and the picture of fine-tuning sliding smoothly along a Wasserstein geodesic does not survive measurement. What we offer is a set of explicit certificates and their conditions - not a free lunch, but a priced one.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.