acceptodds
Under review as a conference paper at ICLR 2027

SEAR: Benchmarking Self-Evolving Agentic Robotics

Abstract

Robots deployed in the real world need to improve from their own experience. We introduce SEAR (Self-Evolving Agentic Robotics), a framework and benchmark for robotic agents that get better with use, whether by rewriting their code, skills, and memories or by training a vision-language-action (VLA) model. Evaluating self-improvement is harder than evaluating a fixed policy, and SEAR addresses five specific problems. First, an agent that tries more often may look like it is learning when it is only sampling more, so we compare against best-of-N baselines with the same budget. Second, getting better at a practiced task is not the same as gen- eralizing, so we test on a five-level ladder: held-out initial states, unseen tasks, a second physics engine, and real hardware. Third, agents can exploit privileged simulator state, so success is scored by a judge the agent cannot access. Fourth, results can depend on the path an agent happens to take, so we run several inde- pendent copies of each agent from the same starting point. Fifth, code-writing and model-training agents spend very different resources, so we track compute along- side wall-clock time. We evaluate on five families of simulated tasks and five physical tasks. Inherited skills raise success from 61.9% to 72.3% on the hardest simulated family, yet inherited mistakes can persist across later tasks. A distilled VLA succeeds under camera faults that disable its own teacher. On hardware, the strongest simulated agent finishes last, and switching from frozen code to live ex- ecution reverses the ranking of the two models tested in both modes. SEAR is the first benchmark that evaluates self-evolving agents on robot tasks under rigorous controls, and we hope it gives future work on this direction a common playground and shapes the community’s understanding of this area.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.