Geometric Regularization of Learned Behavior Spaces in Unsupervised Quality-Diversity Optimization
Abstract
Discovering increasingly capable behaviors can require preserving solutions whose value only becomes apparent when later discoveries build on them. A central challenge is thus retaining these stepping stones without specifying in advance which behavioral differences will matter for future progress. Quality-diversity algorithms pursue this goal by maintaining a repertoire of distinct policies that the search can revisit and improve. For example, learning a latent representation of behavioral attributes directly from actions of a robot removes the need to specify the behavioral dimensions of any given policy that distinguish meaningfully different solutions from one another. Prior works, like AURORA, train autoencoders on action trajectories to find these important dimensions of variation. But autoencoder reconstruction loss alone does not uniquely determine the latent-space distances between potential solutions. Here, we introduce AURORA-VC, which adds variance and covariance penalties to online trajectory reconstruction. These penalties encourage variation across descriptor dimensions and discourage correlations between behavioral descriptor axes independently of the selected loss or fitness functions used later to assess solution quality. With both PGA-AURORA and AURORA, maximum fitness improves in four of ten tasks at an uncorrected paired , with task-dependent effects on repertoire quality. VC is associated with longer survival of high-fitness admissions and reuse of older parents. Survivor-selection interventions support selection as one contributor to the fitness gains. Together, these results show that descriptor learning shapes the search itself, determining which policies compete, survive, and remain available as stepping stones toward higher fitness.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.