acceptodds
Under review as a conference paper at ICLR 2027

KestRL: A Dual-Route JAX Library for Modular Reinforcement Learning

Abstract

We present KestRL, a dual-route reinforcement learning library built on Flax NNX, designed to let researchers build algorithms from reusable components and run them with both host-driven and JAX-native environments. KestRL builds on NNX graph/state separation to organize learning updates around explicit training state. It provides two execution routes: a Standard route, which combines Gymnasium environments with compiled learner updates, and a Compiled route, which brings Brax environments and learning updates into compiled training loops and supports batched execution of independent runs. SAC and PPO illustrate how learning code can be shared across these routes as collection and scheduling adapt to each environment. Controlled update comparisons and complete training experiments examine the resulting execution costs. For host-dispatched updates, the controlled comparisons show improvements over naive NNX dispatch, alongside workloads where cached NNX or hand-written functional JAX remains faster. In a five-run Ant PPO study at a fixed environment-step budget, KestRL Standard reduces end-to-end time by 32.8% relative to eager CleanRL and 25.7% relative to the evaluated Inductor-enabled configuration. In the reported Brax SAC configurations, KestRL Compiled shows similar independent-run scaling to Rejax, with lower observed amortized runtime on Ant and Hopper across the tested batch sizes.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.