BodyBound: A Benchmark for Diagnosing Robot-Policy Responses to Body Changes Beyond Task Success
Abstract
Task success does not by itself show whether a robot policy adapts when its body changes. A policy may fail before the intervention, or the same action sequence may work on both bodies. BodyBound is a controlled benchmark for separating these cases. It pairs executions with the same scene, goal, initial state, and rollout budget while changing one body property: joint limits, jaw capacity, finger-force limit, or wrist-guard geometry within a Panda topology. An independent trace checker is combined with four diagnostics: familiar-start competence, expert attainability witnesses, reciprocal body-description interventions, and fixed-action replay. We release 100 simulated RGB-D demonstrations and evaluate HPT and Diffusion Policy on 60 fresh paired scenes (720 matching-description rollouts). On the public test split, success is 4.2% for HPT and 19.4% for Diffusion Policy; all six checkpoints obtain zero successes in Scene OOD. The learned policies also fail on many familiar starts, limiting attribution of failures to body changes. In a competent Portal reference, all 16 fresh pairs succeed with zero description benefit and shared successful actions; a Grasp reference shows a positive description benefit despite shared solutions. BodyBound therefore measures competence, physical response, and evidence of adaptation separately.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.