Embodied Skill Bench: Benchmarking How Well Agents with Embodied Skills Perform on Physical Tasks
Abstract
LLM agents are increasingly used as high-level controllers for robots, drawing on a growing body of VLA policies, perception models, and procedural know- how from earlier robot runs. These modules are hard for agents to reuse because they differ in interfaces, embodiment assumptions, and conditions of use, and how much they actually help on physical tasks has not been measured. In this paper, we first package them as embodied skills, which extend agent skills by making embodiment assumptions and conditions of use explicit, and build a repository of 250 action, perception, and experience skills from 1,052 existing resources. On this basis we introduce EMBODIEDSKILLBENCH, which runs an agent in a closed observation-action loop and records success, actions, tokens, and wall-clock time. It supports three modes: native control, skill-only execution, and Agent-Directed Reuse (ADR), a mode we propose in which the agent keeps its native controls and decides for each skill whether to invoke, modify, or ignore it. Across four agents on 100 perturbed LIBERO-Pro tasks, we find that skills help an agent only as far as they exceed what it can already do. Skill-only execution lifts GPT-5.6- Sol from 11% to 63% and Kimi-K3 and Claude-Opus-5 by about 20 points, while cutting their token usage by 58% to 85%. Yet all four agents then end up between 61% and 64%, and GPT-6-Astra falls from 92% under native control to 61%. This suggests that as agents grow stronger, the role of embodied skills shifts from supplying missing capability to reducing cost: GPT-6-Astra already solves most tasks through native control, so restricting it to skills removes more than it adds. Under ADR, which lets it keep native control and call skills selectively, GPT-6-Astra recovers to 85%, within 7 points of native control, while using 14.1% fewer tokens and 11.9% less time. With Jev as a skill choosing motion directions and gripper actions, it reaches 81% with 50% fewer tokens and 57% less time.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.