acceptodds
Under review as a conference paper at ICLR 2027

Reinforcement Learning Enables Scaling Generalist Robot Policy Improvement

Abstract

The introduction of generalist robot policies has enabled a single, large model to learn a variety of tasks from a vast pre-training corpus. While these models often display powerful zero-shot generalization capabilities, they nonetheless typically require an adaptation phase to reach sufficient performance in deployment. Iterated self-improvement via reinforcement learning (RL) has demonstrated promise in enabling this adaptation, yet such approaches have been largely explored in single-task settings – where the goal is to simply improve on a narrow deployment task – which lies in stark contrast to the multi-task capabilities of the base policy. In this work we ask whether generalist robot policies are capable of multi-task, generalist improvement in deployment – targeting many diverse tasks simultaneously – and what algorithmic recipe enables this. To answer this question, we first develop a set of structured evaluations to test generalist improvement capabilities, and systematically evaluate algorithmic and practical design choices regarding value estimation, policy extraction and resource allocation. We show that not only generalist policies can be improved online across diverse tasks, but that the right recipe can enable generalization and increase performance on tasks not considered during online training. We characterize when this generalization occurs, and show that the training task distribution as well as the use of value-based RL are both critical. To the best of our knowledge, this is the first evidence for online generalist improvement for robot foundation models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.