Beyond Decision Agreement: Evaluating Reinforcement Learning Policy Behavior and Outcomes in Power-System Balancing
Abstract
Learned decision-support policies are often evaluated by their agreement with historical operator decisions. However, historical agreement does not reveal how a policy behaves when its decisions are executed sequentially, because each decision changes the state for subsequent decisions. We study this issue in offline reinforcement learning for power-system balancing at a Belgian transmission system operator (TSO). We evaluate learned policies using three complementary measures: , historical agreement with operator decisions; , behavior during closed-loop execution; and , the resulting system-level performance, measured by mean absolute Area Control Error (ACE). Across controlled changes to the reward, supervision, and decision-time deferral, we find that these measures can change independently. In particular, policies with similar historical agreement can produce substantially different closed-loop behavior, while deliberate improvements in agreement also alter closed-loop behavior. Yet, across all experiments, these behavioral differences correspond to only small, statistically undetectable changes in the resulting system imbalance. We conclude that historical decision agreement alone is insufficient to validate a learned decision-support policy. Instead, evaluating policy behavior alongside historical agreement and operational outcomes is essential for offline sequential decision support.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.