acceptodds
Under review as a conference paper at ICLR 2027

DISC: Dataset Distillation for Sub-1 IPC Learning under Implicit Shared Representations with Coresets

Abstract

Distilling high-resolution datasets under tight storage budgets requires synthetic representations that are both compact and useful for downstream learning. We introduce Dataset Distillation for Sub-1 IPC Learning under Implicit Shared Representations with Coresets (DISC), a framework for the sub-1 images-per-class (IPC) regime, where the storage budget is smaller than that required for 1 IPC. Building on the storage efficiency of implicit neural representations (INRs), DISC further reduces per-instance storage costs by sharing a coordinate encoder across synthetic images while retaining lightweight instance-specific decoders. This allows more synthetic instances to be stored within the same byte budget. To translate this additional capacity into effective training data, DISC couples parameter sharing with a two-stage coreset initialization: median-distance samples pretrain the shared encoder on non-outlier variation, and prototype-nearest samples warm up the encoder and decoders with representative examples before distillation. On six ImageNet subsets at a 0.5 IPC-equivalent budget, DISC improves mean Top-1 accuracy over DDiF by 4.6 and 5.4 percentage points at and , respectively, under distribution matching. Further improvements under gradient matching and in average cross-architecture performance demonstrate the effectiveness of combining shared representations with stage-specific initialization for a better accuracy–storage trade-off in highly compressed dataset distillation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.