Plasmax: Differentiable & Parallelizable Environments for Tokamak Control
Abstract
Controlling unstable plasma in tokamaks to improve nuclear fusion performance remains an open problem, and any progress in this area brings us closer to efficient nuclear fusion power for carbon-free energy. While Reinforcement Learning (RL) has great potential to tackle such complex dynamical control problems, research on its application to fusion remains sparse. One significant barrier is the lack of fast, configurable RL environments tailored to fusion applications. In this work, we introduce _Plasmax_: 17 differentiable, parallelizable environments for tokamak kinetic control with 5 physics models for dynamics simulation. Plasmax enables benchmarking RL methods against other open- and closed-loop strategies, evaluating robustness to real-world-inspired observation corruption (e.g., noise or delays), and studying transfer between simulators with different fidelities. We show that Plasmax is significantly faster than prior work, enabling up to a speedup and training an RL agent in 7 minutes on hardware accelerators. We benchmark a wide array of trajectory and policy optimization methods, including RL, Evolution Strategies (ES), and differentiating through the environment. We find that open-loop control optimized with CMA-ES performs 32% better than the best RL method we evaluate (SAC). We also investigate how trained policies generalize to different backends, how real-world-inspired observation corruptions affect training, and how exploration can be used as a bug-hunting tool to improve simulators. We hope Plasmax will foster collaboration between machine learning and fusion researchers to develop practical fusion energy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.