How Dynamic Batch Schedules Shape Final Prediction Risk
Abstract
We develop a finite-horizon statistical theory of dynamic mini-batching under warmup–stable–decay learning rates (WSD) for sketched linear regression with power-law spectral and source conditions. WSD induces two effective time scales: deterministic learning is governed by the full horizon , whereas stochastic noise retains only a shorter terminal memory . Although a dynamic batch schedule has one degree of freedom per update, its leading stochastic contribution to final risk collapses to an influence-weighted scalar summary: for one-pass WSD-SGD and under data reuse. This compression also reveals a structural distinction. One-pass batching controls centered SGD variance, while multi-pass batching controls only fluctuation around the full-batch WSD-GD reference; approximation, GD bias, and GD variance remain schedule-independent at a fixed update horizon. We obtain matching one-pass variance bounds and nearly matching multi-pass fluctuation bounds. As a consequence, minimizing the influence-weighted stochastic terms yields the oracle allocation .
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.