LISC: Balancing Sample Influence with Spectral Reweighting
Abstract
Modern datasets are highly heterogeneous, with predictive signals unevenly distributed across samples. As a consequence, gradients estimated as averages over mini-batches tend to reinforce dominant learning signals, leaving other potentially useful directions weakly represented. This issue cannot be directly addressed by optimizers that operate strictly on gradients already aggregated across samples. We therefore propose a sample-reweighting method, Local Inverse Spectral Curriculum (LISC), which re-balances gradient directions while remaining fully compatible with existing optimizers. Starting from a local linear approximation of per-example losses, LISC assigns weights that reduce the relative influence of strongly represented gradient directions. Its underlying spectral transformation is equivalent to preconditioning with a local gradient second moment, yet it avoids constructing a full parameter-space preconditioner. Across vision and language tasks, architectures, and optimizers, LISC improves early-training test performance, with gains exceeding five percentage points in multiple settings, while matching final performance in vision and improving final generalization in several language tasks. These findings indicate that geometry-aware sample aggregation can complement optimizer-level transformations to accelerate learning without sacrificing, and in some settings even improving, final generalization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.