acceptodds
Under review as a conference paper at ICLR 2027

Differentiable evolution of agile agents

Abstract

It is usually much easier to compute meaningful policy gradients than meaningful design gradients, especially in jointed rigid bodies that move quickly and gracefully on land. Whereas soft bodies without bones and joints naturally lend themselves to continuous modeling, they are generally slower, weaker, and less interesting in terrestrial environments. Here we demonstrate end-to-end differentiable evolution of articulated agents capable of agile legged locomotion. Starting from a primordial capsule, jointed links emerge and grow and move along the body as others shrink and disappear, across evolutionary time, depending on derivatives of fitness with respect to design variables. These variables were sampled from a distribution to produce a population of designs, which were controlled by a shared universal policy. Gradients were averaged across individuals in the population and across truncated learning windows inside each forward simulation to update the design distribution and universal controller concurrently. This enables the sample-efficient discovery and gradient-based optimization of novel robots with exciting athletic behaviors.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.