A Unified Multi-task Benchmark for Data-driven Models of Power Systems
Abstract
Can a learned representation of a structured physical system transfer across tasks, topologies, scales, and interventions? Despite growing interest in foundation models for scientific and engineering systems, existing power-grid datasets largely isolate individual tasks and operating regimes, making such questions difficult to study. We introduce a unified multi-task benchmark for power systems whose tasks are genuinely different learning problems: constrained regression (AC and distribution optimal power flow), mixed-integer structured prediction (security-constrained unit commitment), rare-event screening ( contingency analysis), inverse problems under partial observability (transmission and distribution state estimation), and Monte-Carlo risk estimation (reliability assessment), spanning transmission and distribution networks ranging from 3 to 13659 buses. Alongside the data, we specify benchmark results on cross-task and topology transfer. Surprisingly, stock GNNs, cross-task pretraining, and publicly released foundation models frequently fail to outperform simple classical or nearest-neighbor baselines, revealing substantial headroom for models that truly generalize across structured physical systems. We release the dataset, generalization splits, and evaluation harness as an open benchmark for graph learning, scientific foundation models, and structured prediction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.