acceptodds
Under review as a conference paper at ICLR 2027

Recommender regret for improving regularised RL estimation with bandits and MCTS

Abstract

Regularised reinforcement learning (RL) algorithms rely on estimators, like Monte Carlo tree search or one-step look-ahead, to provide policy, value, or Q-value estimates in training and inference. However, existing bandit-based estimation frameworks evaluate sampling efficiency, not the quality of estimates. To evaluate the quality of a recommendation (estimate), we define regulariser-aware recommender regrets, analysing their lower bounds, and proving that simple algorithms like uniform sampling can achieve these bounds. Finally, we unify bandit- and search-based target improvement under Recommender MCTS. Across synthetic trees, arcade games, and scientific discovery tasks, recommenders yield faster target convergence and superior downstream performance. Our experiments characterise for the first time the potential of regularised MCTS in scientific applications.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.