acceptodds
Under review as a conference paper at ICLR 2027

What a Robot Harness Asks Its VLM: A Deployment-Grounded Benchmark

Abstract

The release of GPT-6 Astra has intensified interest in embodied systems built around vision–language models (VLMs) and robot-specific harnesses. In such systems, however, VLMs are still commonly evaluated as standalone models rather than through the interfaces at which they operate. Even existing benchmarks grounded in real-world embodied settings typically measure either end-to-end system performance or isolated capabilities such as spatial reasoning, abstracting away the interface through which a robot software stack queries its VLM. Their results therefore provide limited guidance for model selection and diagnosis in deployed embodied systems. We address this gap by introducing and releasing a deployment-grounded benchmark derived from requests logged during the routine operation of a manipulation robot already deployed in a real-world environment. Rather than constructing tasks for evaluation, we derive both the task taxonomy and the evaluation instances from actual calls issued by the robot stack. The benchmark covers eight request categories and preserves the task-relevant processed views and output contracts used in deployment. It contains category-balanced, semantically aligned Chinese–English question pairs. We evaluate representative VLMs with output-contract-specific metrics that distinguish format violations from semantic errors. This setup enables a deployment-grounded comparison of VLMs at the robot–VLM interface through which they are actually used.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.