acceptodds
Under review as a conference paper at ICLR 2027

ALOHA: Agentic LLM-driven Online Hyperparameter Adaptation for Reinforcement Learning

Abstract

Hyperparameter optimization (HPO) for reinforcement learning is computationally expensive. Offline methods such as Bayesian optimization require hundreds of complete training runs while population-based methods demand dozens of parallel agents, making both approaches impractical when simulation costs are high relative to the available compute budget. We introduce ALOHA (Agentic LLM-driven Online Hyperparameter Adaptation), a method for online hyperparameter optimization within a single training run, that embeds a large language model directly into the RL training loop to adaptively adjust hyperparameters at fixed intervention points based on online training metrics, requiring no preprocessing and no parallel agents. At each intervention, ALOHA produces updated hyperparameter values alongside natural-language reasoning traces that explicitly justify each adjustment, making each decision interpretable. We evaluate ALOHA across three RL algorithms (PPO, DQN, SAC) and a diverse suite of environments spanning continuous locomotion, discrete control, and grid navigation, benchmarked using ARLBench. Our results demonstrate that ALOHA matches or exceeds Bayesian optimization in 9 of 14 environments and outperforms population-based training across all 14, requiring no offline search budget compared to Bayesian optimization SMAC's 400 total training runs (80 trials 5 seeds) per environment, without any preprocessing runs, parallel agents, or domain-specific tuning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.