Physical Agency: What the Interface Unlocks
Abstract
Frontier models increasingly act as robot-use agents, calling a robot's motor skills as tools much as they interact with a browser or code editor. But exposing skills as tools leaves open how the model should use them: what it emits, what feedback it receives, and when it decides again. We study the interface through controlled comparisons with the robot, cameras, scenes, and motor-policy weights held fixed, and separately measure the benefit of adding a complementary frozen skill. We introduce Pigey (Physical Agency), an agent that (1) issues language-level actions one at a time over a persistent task state, (2) uses verified execution, checking post-action images and robot state and recovering when it detects failure, and (3) draws on an open set of skills exposed as tools it selects for each subgoal. No additional policy training or demonstrations are required. On 30 real-robot tasks designed to probe decisions over existing motor skills, the same frozen π0.5 policy succeeds in 16.7% of trials when called directly and 84.0% through Pigey. Adding a frozen motion planner as a second skill raises success to 97.3%. Removing verification and recovery lowers success by 28 percentage points, the largest ablation drop. Gains also hold in simulation and across Claude, Gemini, and Qwen reasoners. Most of what the policy was missing was an interface, not a skill.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.