Not All Forgetting is Equal: Objective Structure in Temporal Continual Learning
Abstract
Catastrophic forgetting is often treated as a uniform cost of sequential updates, yet on structured temporal streams its severity can depend on the prediction objective. We study Domain-IL fine-tuning on three matched univariate corpora—glucose, electricity, and weather—under five objectives (next-value and short-/long-horizon regression; dense and rare within-horizon events) and five architectures. Event objectives forget far more than value objectives in every domain-architecture cell. Controls rule out rarity, horizon, and the monthly protocol; a reference swap shows the gap follows how the target is defined (entity-fixed thresholds vs. segment-relative value targets), not events as such. Joint-vs-sequential and recurrence diagnostics then split failures into preventable interference and a recurrence gap invisible to single-transition forgetting. Savings, recalibration, and mechanistic probes further show that the loss is often recoverable for recurrent learners, while transformer damage reaches the representation. We demonstrate that not all forgetting on streams is equal. Rather, it is objective-conditioned and diagnostically typed.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.