When Does Training Order Buy Compute? Exact Boundaries and Robust Gain Guarantees
Abstract
When does changing data weights outperform the best fixed mixture, rather than merely a poorly chosen baseline? We give an exact answer for shared-optimum quadratic learning. For two mirror sources at angle , with orthogonal initial-error mass , balanced mixing is optimal over all measurable schedules at budget exactly when . Below the boundary, two mixed phases are enough to improve performance; sufficiently close to it they attain the global gain ceiling exactly. We then distinguish exactness from robustness. Aligned target anisotropy shifts the exact threshold, but arbitrarily small cross moments or unequal source strengths can destroy exact static optimality. A completed-square bound nevertheless limits the reoptimized gain quadratically in the cross moment, while a uniform perturbation bound covers general changes. Finite-budget plots show where candidate schedules approach the all-schedule ceiling and where a gap remains. Exact Gaussian stochastic gradient descent (SGD) experiments are indexed by distance from the boundary and control source composition through stratified batches and rapid alternation. A uniform SGD-to-flow bound guarantees convergence, not the observed near-boundary gain rate. The results identify both the value and the fragility of training order, without claiming a universal neural-network curriculum.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.