RoboARCade: Benchmarking Agentic Robot Control
Abstract
Improvements in multi-modal capabilities have enabled general agents to act directly as robot policies, a setting we refer to as agentic robot control (ARC). We introduce RoboARCade, a benchmark of long-horizon manipulation tasks across multiple embodiments. It is designed to test in-context adaptation, memory, and reasoning for precise control, extending beyond the capabilities emphasized by prior robotic policy benchmarks. The benchmark differentiates leading agents: GPT-6 Astra achieves 49.2% compared with 41.5% for Claude Fable 5.1, leaving substantial room for improvement. We equip agents with different interfaces to study their effects on performance and behavior. Our benchmark identifies several bottlenecks that prevent agents from zero-shot completion of robotic tasks: identifying an approach, executing it precisely, and recognizing subtask completion. Additionally, we find that performance can be improved without additional human input through retries, successful demonstrations, and prompt optimization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.