ManipScope: A Diagnostic Benchmark for Instruction Following and Generalization in Robotic Manipulation
Abstract
Completing a manipulation task does not establish that a robot follows the requested goal rather than a familiar routine. Our proposed ManipScope is an evaluation-only benchmark for robotic instruction following under changed task requirements. Built using RoboTwin 2.0’s simulation resources, it specifies 48 scenes across six task families with explicit source-task links. We deliberately provide no expert training demonstrations for these scenes, targeting skill reuse under changed requests rather than fitting supplied examples. Task-success and selection metrics measure completion and satisfaction of the requested properties. Controlled instruction and layout changes test whether choices track the request, with future-video and internal-intervention probes where access permits. Evaluations of vision–language–action policies, world-action policies, and agents reveal uneven task performance, partial instruction satisfaction, and correct selection without completion. In controlled checkpoint case studies, individual-trial success can coexist with weak paired instruction correctness; forecasts and execution can agree on the wrong target; and internal interventions change predicted actions without establishing useful goal transfer. These findings show why success, consistency, and sensitivity must each be checked against the requested goal.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.