Z-RoboSTAR: Building and Benchmarking Robotic System 2 Agents for Reasoning in Sustained Long-Horizon Manipulation
Abstract
Sustainable long-horizon robotic manipulation requires a deliberative System 2 agent that can reason, plan, and sustain progress toward a goal as the world evolves. Complex multi-step manipulation changes the world state, while interventions may invalidate prior beliefs, requiring System 2 to continually track progress from partial observations. We present Z-RoboSTAR, a harness that enables vision-language models (VLMs) to act as System 2 agents, centered on a dynamically updated symbolic scene graph that serves as a persistent shared-state interface to synchronize scene understanding, task progress, and planning, for failure recovery and resolving reconciliation. To efficiently compare open- and closed-source VLMs for our System 2 harness, we construct Z-RoboSTAR-Bench, an offline open-loop benchmark and evaluation engine spanning 57 tabletop scenes with diverse spatiotemporal dependencies and interventions. The model ranking is preserved under human-executed closed-loop evaluation, supporting the benchmark's relevance to real-world agent performance. Beyond evaluation, we develop Z-RoboSTAR-27B through trajectory-level distillation to improve its robustness over prolonged execution. When deployed on a Galaxea R1 Lite with a fine-tuned System 1 controller, Z-RoboSTAR-27B achieves a 66.9% task success rate on our real-world long-horizon manipulation tasks, ranking second only to GPT-5.5. These results demonstrate effective System 2 coordination of complex, long-horizon manipulation in real-world scenes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.