acceptodds
Under review as a conference paper at ICLR 2027

Asynchronous Sim-Real Co-training for On-Policy Robot Learning with Limited Real Hardware

Abstract

Improving an already capable robot policy through online interaction is challenging when physical hardware and interaction opportunities are limited. Simulation can supply additional experience, but differences in collection speed and availability across domains complicate batch composition and data freshness during joint online learning. We present Asynchronous Sim-Real Co-training (ASRC), a framework that uses continually collected simulated and real-world experience for PPO-style policy improvement. ASRC decouples collection from policy optimization, allowing independent interaction in each domain while learning proceeds. To coordinate their heterogeneous experience, an availability-aware scheduler adjusts the real-data fraction within a prescribed range, while latest-first selection and bounded-buffer eviction prioritize recent samples. The framework supports both CNN policy optimization and residual adaptation of . Across three manipulation tasks on a single physical robot, both policy families achieve final real-world success rates above 90% on every task, outperforming the evaluated baselines. System experiments show up to higher update throughput than a synchronous counterpart, while ablations support the importance of balancing domain composition and prioritizing fresh experience.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.