acceptodds
Under review as a conference paper at ICLR 2027

Robo-Harness K1: Harnessing Robot-Use Agent via Perception Augmentation

Abstract

Foundation vision-language models (VLMs) recognize objects, interpret instructions, and reason about spatial relations. A central question for robotic manipulation is how to fully translate this capability into physical actions. Adding a learned action head, as in vision-language-action (VLA) models, requires large-scale demonstrations, may degrade the inherent VL understanding of VLMs, and generalizes poorly. Recent results of Astra (GPT-6) as policy show the potential of robot-use agents (RUAs), which keep the VLM intact as a direct controller. However, this route is slow, expensive, and difficult for weaker models. We step back and ask: what is fundamentally missing between a capable VLM and a working manipulation policy? We find that the answer is accessible perceptual toolkits for 3D spatial understanding and measurement. We introduce Robo-Harness K1, which exposes perception as tools for harness: the agent queries calibrated depth, inspects visual anchors with persistent object identities, evaluates grasp hypotheses, and makes actions and motions from the returned evidence. Robo-Harness K1 enables the VLM to understand depth and space without invasive depth-encoder retraining, so the agent can reason about 3D object relations instead of guessing. On LIBERO-PRO, Gemini 3.7 Flash with K1 reaches 77.8%, surpassing GPT-6 Astra's 61.1% with an RGB-only harness, and K1 further augments Astra to 88.9%. Without target fine-tuning, Gemini with K1 transfers to three RoboSuite arms (90.0% average on shared tasks) and dual-arm RoboTwin tasks. For training RUA models, K1 provides better sample efficiency and generalization than VLAs and vanilla RUA. A Qwen3.5-9B model fine-tuned on only 107 episodes reaches 44.2% on new-state generalization versus 30.2% for OpenVLA, and 13.9% on new-task generalization versus 0.0% for OpenVLA. These results suggest that training RUAs with a proper perceptual harness outperforms VLAs, unlocking VLMs' great potential for robotic manipulation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.