Patch-Level Codebook Parameterization for Dataset Distillation
Abstract
Dataset distillation aims to synthesize a compact dataset that retains the essential knowledge of a large-scale dataset, enabling efficient model training while significantly reducing storage and computational costs. Recent parameterization methods enhance storage efficiency through compact or shared representations, but recurring local structures may still be redundantly encoded across synthetic images, limiting storage efficiency. In this paper, we propose to parameterize a synthetic dataset using patch-level learnable codebooks and patch index maps, such that each synthetic image is generated by assembling learned representative patches according to the index map. By explicitly sharing patch-level representations across synthetic images, our method reduces repeated encoding of recurring local structures and improves the use of limited storage budgets. Furthermore, we employ multiple codebooks at different stages and combine their patches additively to further improve representation power. Extensive experiments on standard benchmarks demonstrate that our method achieves state-of-the-art performance across various dataset distillation methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.