EmbodiedSWE: Coding Agents for Long Horizon Dexterous Robotics
Abstract
We study coding agents for long-horizon, dexterous robotics and ask whether their solutions can provide scalable supervision for learning general robot policies. To test this, we develop EmbodiedSWE-Bench, an agent-native simulation benchmark spanning contact-rich manipulation, deformable objects, and long-horizon tasks requiring up to half an hour of continuous interaction. We find that frontier coding agents can solve complex long-horizon tasks and transfer prior solutions across both tasks and embodiments. We also design supporting tools that help agents more effectively solve these tasks. However, the resulting solutions require substantial iterative interaction and are typically specialized to individual task instances. We therefore introduce EmbodiedSWE-Gen, which expands a single solution from coding agent into large diverse trajectories for training a VLA. Finally, we present a pipeline that automatically generates diverse, verifiable simulation environments and explore improving coding agents through reinforcement learning. Together, our framework casts coding agents as solvers, teachers, and students in a scalable robot-learning loop.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.