acceptodds
Under review as a conference paper at ICLR 2027

Building Exploratory Robot Policies via Adversarial Mid Training

Abstract

Robot policies today are predominantly trained on human-collected data, which naturally concentrates on competent, successful behavior. As a result, such data provides limited coverage of off-distribution states and diverse low-level behaviors needed to solve complex tasks over long horizons or improve through their own practice. In this paper, we study how to explicitly build this coverage into general-purpose robot policies. Specifically, we study a distinct collection protocol, adversarial data collection: while a human attempts a task, an adversary reversibly perturbs the environment or the task, forcing the operator to encounter unfamiliar states and adapt via low-level behaviors. We collect such rollouts, relabel segments with natural-language annotations identifying the local outcome pursued in each segment, and use them to midtrain a history-conditioned pretrained robot policy. The resulting policy retries after failures, makes sustained progress over longer horizons, and internalizes diverse ways of solving a given task. Across several long-horizon tasks on three real-robot platforms and in simulation, compared to training with standard demonstration data, our approach yields better scaling of zero-shot performance, greater robustness and generalization to novel configurations, gains from additional test-time compute (i.e., longer rollouts), and faster improvement during online finetuning. Project page and robot videos are available at https://explr-beta.pages.dev/https://explr-beta.pages.dev/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.