The Grid as Observable State: Infrastructure-Aware Resilience for Large-Scale Training
Abstract
Large-scale distributed training relies on checkpointing to recover from failures. Its resilience also depends on the infrastructure shared by the sites that host it. We argue that resilience evaluations should make this shared exposure visible alongside hardware failure statistics. Using five years of hourly ERCOT electricity-price data, we characterize the spatial and temporal structure of negative-price episodes. In 2025, negative-price indicators at the studied North Texas substations have pairwise correlations of 0.875–0.933. In the corridor-level series, episodes last 3.8 hours on average, and 46 of 56 episode starts (82.1%) fall between 03:00 and 12:00 CST, a window covering 37.5% of the day. These observations reveal shared market exposure and concentrated timing, two dimensions of infrastructure context relevant to resilience evaluation. Whether market conditions affect training availability depends on the facility's operational response. Under an explicit episode-to-interruption scenario, a checkpointing illustration connects the observed timing to a break-even criterion involving interruption exposure, checkpoint effort, and recovery costs. Comparisons at matched nominal checkpoint frequency show why resource accounting matters when assessing adaptation. Together, the empirical findings and conditional analysis motivate evaluating observable infrastructure signals alongside facility telemetry to identify when they can improve training-resilience decisions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.