SatelliteOpsGym: Benchmarking AI Agents from Fault Diagnosis to Closed-Loop Recovery in Satellite Operations}
Abstract
Autonomous satellite operations require intelligent agents to diagnose faults, execute interventions, and verify recovery through closed-loop interaction. However, the operational capabilities and failure boundaries of existing agents remain poorly understood. We introduce SatelliteOpsGym, an interactive satellite operations environment for systematically evaluating agents and supporting their future training and improvement. Built upon documented on-orbit faults, real multichannel telemetry, and causal fault graphs, the environment enables agents to investigate anomalies, execute constrained commands, observe their consequences, and verify recovery. Its operational fidelity is assessed through double-blind evaluation by certified satellite engineers. Experiments with eight frontier LLMs on 15 shared scenarios reveal a pronounced fault-to-recovery gap: although fault identification reaches 71.4% and command validity reaches 100%, root-cause Planning and Recovery achieve at most 39.3% and 32.1%, respectively. By exposing failures across the operational workflow, SatelliteOpsGym provides insights for developing more reliable satellite operations agents.Data and code are available at https://github.com/anonymous-satelliteopsgym/SatelliteOpsGym.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.