acceptodds
Under review as a conference paper at ICLR 2027

DexPolicyEvo: Evolution of Dexterous Grasping Policies through Failure-Driven Self-Improvement

Abstract

Enabling dexterous grasping policies to continually improve toward open-vocabulary object grasping remains an open challenge. Achieving this efficiently requires understanding a policy's current capabilities and weaknesses to tailor its training accordingly. With this goal in mind, we introduce DexPolicyEvo, a system that enables dexterous grasping policies to evolve autonomously through failure-driven self-improvement. Its key idea is to connect failure diagnosis with targeted training-data construction and geometric curriculum design through measurable and editable object properties. Six collaborating agents implement this improvement loop: the Diagnostic Designer and Policy Evaluator construct tests across everyday scenarios and assess policy capabilities; the Failure Analyst and Repair Designer identify weaknesses and devise targeted training data and curricula; and the Policy Trainer and Policy Verifier update the policy and validate its improvement. As the policy improves, the agents revise its training plan to address emerging weaknesses and further expand its grasping capabilities. Experiments demonstrate that DexPolicyEvo continually improves policy generalization while preserving existing grasping capabilities.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.