Dual-Agent Co-Design of Terrain Curricula and Reward Functions for Mixed-Terrain Quadrupedal Locomotion
Abstract
Robust quadrupedal locomotion over mixed terrain requires the robot to preserve stable motion as contact, posture, and momentum demands change between ter- rain segments. These transitions couple two training decisions: the terrain cur- ricula used for policy training and the reward functions used to guide policy op- timization. Optimizing them separately can misalign reward incentives with the behaviors required by the training terrains. We introduce a policy-mediated dual- agent framework in which a Terrain Curriculum Agent and a Reward Hypothe- sis Agent co-design executable mixed-terrain training arenas and arena-informed reward programs. Transition-frontier feedback computed from rollouts of the in- duced policy identifies unresolved transitions and links them to local geometry and robot–terrain interactions, enabling both agents to address the same locomo- tion bottleneck. Candidate designs are grounded by proximal policy optimization in physics simulation and selected through hierarchical behavioral validation that combines reward screening with gait-constrained cross-arena evaluation. Exper- iments on an independently constructed mixed-terrain test set show gains over terrain-evolution, reward-search, hand-designed, and task-curriculum baselines in full-course success and traversal progress. The advantage is larger on closely spaced terrain compositions than on isolated components, and ablations show that terrain evolution and explicit designer feedback complement reward optimization, particularly on difficult courses. Additional evaluations on three public terrain suites support the transfer of the learned policies across terrain generators.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.