Renewable-Aware Neural Network Training with Phase-Dependent Lyapunov Control
Abstract
In neural network training on GPUs, a larger batch shortens an epoch but draws more power, and its effect on the loss changes as training progresses. When renewable supply varies, the best batch size therefore depends on the available power and the training phase. Since the batch size changes how long each training segment takes, we formulate renewable-aware training as a renewal control problem that minimizes the renewable shortfall per unit time subject to a duration-weighted error constraint. We enforce the error budget with a virtual queue and design a drift-plus-penalty ratio controller that penalizes its backlog with phase-dependent Lyapunov weights. We prove an bound on the shortfall-rate gap with an average virtual backlog and extend it to finite Markov states and recurrent training classes. To run the controller in GPU training, we score each candidate action with a predictor calibrated from shared-checkpoint measurements and update the queue and the phase with the measured loss. We evaluate the controller in 75 training runs with independent supply and a measured photovoltaic (PV) day. Under independent supply, the controller reaches a common quality target in 59.7% less time than batch 16 at a similar shortfall rate, and relative to batch 128 it reduces the shortfall rate by 72.5% with a 13.4% longer target time. With a noon start on the PV day, its shortfall rate is 32.7% lower than that of batch 128 at the same target time, and as the PV capacity grows, its DWAE falls from 0.040 to 0.026 while fixed schedules keep their error. Emphasizing the backlog during rapid descent shortens the mean target time by 7.0% relative to equal weights.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.