acceptodds
Under review as a conference paper at ICLR 2027

The Cost of Hyperparameter Tuning in Online Reinforcement Learning

Abstract

The performance of reinforcement learning (RL) algorithms is often benchmarked without accounting for the cost of hyperparameter tuning, despite its significant practical impact. In this paper, we show that such practices can misrepresent the perceived efficiency of RL methods and impede meaningful algorithmic progress. We formalize this concern by proving a lower bound showing that tuning hyperparameters in RL can induce an exponential blow-up in the sample complexity or regret, in stark contrast to the linear overhead observed in supervised learning. This highlights a fundamental inefficiency specific to the online learning setting. In light of this, we propose evaluation protocols that account for the number and cost of tuned hyperparameters, enabling fairer comparisons across algorithms. Surprisingly, we find that once tuning cost is accounted for, elementary algorithms can outperform their successors with more sophisticated design. These findings suggest the value of a tuning-aware evaluation perspective in how RL algorithms are benchmarked and compared, especially in settings where efficiency and scalability are critical.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.