acceptodds
Under review as a conference paper at ICLR 2027

From Local Condensation to Global Utility: Replay-Enhanced Continual Dataset Distillation

Abstract

Dataset distillation typically assumes that the entire training dataset is available simultaneously. In this work, we study *Continual Dataset Distillation* (CDD), a new setting in which class groups arrive sequentially and a synthetic memory must be constructed and maintained without retaining historical real samples. The resulting synthetic memory is evaluated by training fresh models on all observed classes, requiring synthetic subsets distilled at different steps to remain compatible as the label space expands. We identify a key failure mode in CDD, which we term *Local-Global Discrimination Collapse* (LGDC). Synthetic subsets distilled independently within individual class groups may remain effective for local classification, yet lose compatibility after aggregation for global classification. To address this issue, we propose *Replay-Enhanced Continual Dataset Distillation* (ReCDD), which replays historical synthetic samples during the distillation of new classes and refines the historical memory as the label space expands. To provide a consistent competition signal when future classes are unavailable, ReCDD introduces a non-semantic auxiliary reference that is discarded before evaluation. Experiments on four datasets and five representative methods spanning gradient and trajectory matching show that ReCDD improves final global accuracy in most evaluated settings. Further diagnostic analyses show that ReCDD reduces the local-global discrepancy and cross-step errors associated with LGDC.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.