GroundProbe: A Benchmark for Diagnosing Language Grounding in Robotic Policies
Abstract
Vision-Language-Action (VLA) models report high success rates on standard manipulation benchmarks, yet recent stress tests reveal that this success often collapses under modest perturbations of objects, viewpoints, or language, suggesting that headline numbers conflate trajectory memorization with functional understanding. We introduce benchmark, a diagnostic benchmark that decomposes policy evaluation into three complementary dimensions, namely (1) language grounding, (2) compositional generalization, and (3) state-conditioned behavior. Within each dimension, training and testing splits are constructed so that the building blocks for the testing distribution are present in training, and each testing condition is matched to a seen one that differs in a single controlled factor, so that, given competence on the matched seen condition, a drop on the held-out one is evidence of a failure to transfer rather than of insufficient supervision. benchmark is built on NVIDIA Isaac Lab-Arena with a Franka Panda manipulator and a curated set of 30 meshes, totaling roughly 2,500 training trajectories and 1,770 testing episodes spread across six sub-tasks. Rather than reducing a policy to one number, benchmark reports closed-loop success rates together with normalized transfer gaps, a gap being the fall from a seen composition to the held-out one that differs from it in a single controlled factor, which distinguishes failure modes such as direction being entangled with category from the absence of an abstract category representation. The release comprises the scenes, instruction sets, teleoperated demonstrations, evaluation configurations and the analysis code, giving VLA designers a sharper diagnostic instrument than overall task success rate. We report a first datapoint rather than a full diagnostic profile. Populating the three axes across a slate of current models is the work this release is meant to enable. The benchmark, the demonstrations and the analysis code will be released upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.