acceptodds
Under review as a conference paper at ICLR 2027

In-MANU: Probing the Capabilities and Self-Improvement of Frontier Agents in Robotic Manipulation

Abstract

Frontier multimodal models can act as robot policies without robot-specific training: they interpret camera images and task instructions and issue end-effector commands directly. The extent of this capability relative to post-trained embodied foundation models, and the extent to which such an agent can improve it from its own experience, have not been characterized. We present In-MANU, a study of frontier agents in robotic manipulation that evaluates GPT-6-Astra on the 26 test tasks of the EBench mobile manipulation benchmark and compares it with seven post-trained vision–language–action and world-action policies. With a single annotated demonstration per task, the agent ranks second among the eight systems. It attains the highest success rate on short-horizon tasks and exhibits exploration, recovery from disrupted subgoals and correction after failed attempts within an episode, but it ranks below most post-trained policies on long-horizon and tabletop dexterous-and-precise tasks, which require precise alignment, physical contact and bimanual coordination. We further examine two means of reducing this gap. Access to a post-trained policy as a callable tool increases success only on tasks that the policy solves on its own, and the agent's delegation does not correspond to the competence of the policy. A self-improvement loop, in which a critic instantiated from the same model revises the agent's playbook and harness on the basis of execution feedback, alters the agent's strategy without a statistically significant improvement in test success, both for GPT-6-Astra and, in a smaller replication, for Claude Opus 5.5; the critic's episode-level motion-error reports coincide with measured deviations, but its aggregate account attributes all failure modes to execution.These findings suggest that physical execution reliability remains a major bottleneck for frontier agents under the evaluated control interface.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.