Dynamic Sketching for Efficient SGD via Kernel Mean Approximation with Random Features
Abstract
Large-scale deep learning is often trained on highly redundant datasets, where many examples provide similar learning signals. This raises a natural question: can we reduce training compute by replacing full-data SGD with a small coreset that adapts as the model representation evolves? We propose DySk (Dynamic Sketching), a weighted dynamic coreset method based on loss-aware kernel mean approximation. At periodic refresh steps, DySk computes current representations and losses, builds a loss-weighted random Fourier feature sketch of the representation distribution, and selects a compact support by greedy residual matching. It then fits nonnegative simplex weights on the selected support and trains on the resulting weighted coreset until the next refresh. This yields a forward-pass-only refresh mechanism that tracks the evolving training distribution without requiring per-example backward gradients for selection. Empirically, DySk preserves accuracy while substantially reducing the amount of data processed by SGD, yielding a favorable accuracy–runtime trade-off against gradient-matching and other coreset baselines. We further provide ablations and a conditional sketch-to-gradient analysis supporting the empirical and theoretical behavior of DySk.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.