acceptodds
Under review as a conference paper at ICLR 2027

Physics-Aware Distillation of a Grid Foundation Model for Real-Time Edge Inference

Abstract

Grid foundation models (GridFMs) are graph transformers pretrained on a masked power-flow objective, and once trained they evaluate power flow orders of magnitude faster than an iterative solver. Their size, however, targets data-centre inference, whereas low-inertia operation - grids with little rotating reserve, where disturbances evolve within milliseconds - increasingly asks for inference at the edge, next to the measurement, at batch size one and within a single synchrophasor (PMU) reporting frame. We distil a 20.1M-parameter, 12-layer GridFM teacher into a 0.62M-parameter, 6-layer student, a 32.5× parameter reduction at half the message-passing depth, using an objective that combines the masked powerflow loss with output, physics-residual, latent-state and attention transfer. The physics-residual term, which asks the student to reproduce the teacher’s active- and reactive-power residual profile rather than only to minimise its own, has not previously been applied to typed power-grid graphs. We evaluate on a renewable variant of the IEEE 118-bus benchmark grid in which 26 of the conventional (synchronous) generators are replaced by inverter-based renewables with battery flexibility and standardised reactive-power limits (IEEE 1547), giving 97 529 physically valid (AC-OPF-feasible) operating scenarios, and we ablate all eight combinations of the transfer terms, reporting accuracy against ground truth and fidelity to the teacher separately. Distillation lowers student error on all four predicted quantities, by up to 48% on active-power dispatch, and every distilled student is more accurate than its teacher on reactive-power dispatch. The gain is largely carried by the output matching term. The physics-residual term, in the layer-aggregated form used here, is consistently harmful, which we attribute to the spatial averaging it applies to a bus-level residual field. We study the effect of the parameter reduction on latency on the target device to validate the effectiveness of the reduced layer count.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.