acceptodds
Under review as a conference paper at ICLR 2027

SmartGridBench: Cohort-Bound Evaluation of Industrial LLM Agents for Transformer Maintenance

Abstract

Power-transformer maintenance requires an agent to combine telemetry, fault interpretation, forecasting, and work-order records before recommending an action. We introduce SmartGridBench, a synthetic transformer-maintenance extension of AssetOpsBench, and use it to separate tool-interface choices from planning choices. The frozen resource contains 36 structurally validated scenarios; the principal quality comparison evaluates 31 hand-authored scenarios, while a later repository expansion to 61 scenarios was not evaluated in these experiments. On the 31-scenario core, each configuration has five trials per scenario: direct tool calls receive 67/155 recorded judge passes, MCP Agent-as-Tool receives 57/155, and Verified Plan-Execute receives 86/155. A separate six-scenario latency cohort records median completion times of 8.51 s, 12.22 s, and 37.92 s for those three configurations; the cohorts are not pooled. A 21-case, outcome-stratified cross-family judge check agrees with the historical pass label on 13/21 cases (61.9%) and assigns 4 versus 10 passes, which demonstrates evaluator sensitivity rather than judge accuracy. The synthetic fleet fails an existing DGA realism check, and independent human and domain validation remain unavailable. We therefore present SmartGridBench as a reproducible fixture for conditional systems evaluation; establishing maintenance readiness or deployment safety would require further validation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.