acceptodds
Under review as a conference paper at ICLR 2027

ManipScope: A Diagnostic Benchmark for Instruction Following and Generalization in Robotic Manipulation

Abstract

Completing a manipulation task does not establish that a robot follows the requested goal rather than a familiar routine. Our proposed ManipScope is an evaluation-only benchmark for robotic instruction following under changed task requirements. Built using RoboTwin 2.0’s simulation resources, it specifies 48 scenes across six task families with explicit source-task links. We deliberately provide no expert training demonstrations for these scenes, targeting skill reuse under changed requests rather than fitting supplied examples. Task-success and selection metrics measure completion and satisfaction of the requested properties. Controlled instruction and layout changes test whether choices track the request, with future-video and internal-intervention probes where access permits. Evaluations of vision–language–action policies, world-action policies, and agents reveal uneven task performance, partial instruction satisfaction, and correct selection without completion. In controlled checkpoint case studies, individual-trial success can coexist with weak paired instruction correctness; forecasts and execution can agree on the wrong target; and internal interventions change predicted actions without establishing useful goal transfer. These findings show why success, consistency, and sensitivity must each be checked against the requested goal.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.