Skill2Real: Agentic Skill Learning for Zero-Shot Sim-to-Real Robot Manipulation
Abstract
Sim-to-real learning commonly adapts policies or representations to domain variation, often directing the learning budget toward task-agnostic low-level representations at the expense of task-level knowledge. Skill2Real builds on cross-domain grounding from foundation vision-language models (VLMs) to acquire validated, reusable skills in simulation through a shared code-based policy interface. In its asymmetric Proposer–Verifier–Governor (PVG) loop, the Proposer acts from public observations and API returns, the Verifier diagnoses outcomes with privileged simulation evidence, and the Governor admits updates supported by validation rollouts. A Cerebellum learns reusable local manipulation skills; a Brain then composes the frozen library into long-horizon programs with planning and recovery. At deployment, this hierarchy is grounded in real observations through the same interface. On unseen LIBERO-Pro Long, frozen LIBERO-90 skills raise success from 2.0% to 56.3% with Astra and from 0.5% to 49.0% with Opus 5. Robosuite success reaches 85.1–89.4% across seven tasks. On a real UR5e, the same agent's mean completion rises from 27.50% to 78.75% across four tasks with Skill2Real, a controlled gain that, with cross-model improvements, shows Skill2Real adds a capability orthogonal to the foundation model by learning transferable task knowledge in simulation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.