acceptodds
Under review as a conference paper at ICLR 2027

HORIZON: Recoverable Physical-Domain Expansion for Policy Generalization

Abstract

In on-policy reinforcement learning, expanding physical randomization changes both the control problem and the trajectories that drive learning. Retracting a failed expansion cannot undo learner updates that erode acquired behavior. We introduce HORIZON, a recoverable curriculum that treats physical-domain expansion as a coupled update of the training distribution and learning state. The method accepts an expansion only when the updated policy satisfies candidate-range feasibility and retains performance under previously accepted conditions. After failure, it restores the accepted ranges and learning state while preserving the failed boundary for refinement. Failed exploration thus informs subsequent proposals without its learner updates persisting. In quadruped locomotion, HORIZON achieves 36.8% success under the strongest joint physical shift, compared with 19.8% for OpenAI ADR and 22.7% for GRAM. We further find that broader physical diversity need not improve joint generalization. Under the evaluated training allocations, a four-group curriculum achieves a higher mean success on the tested joint shifts than full seven-group expansion, while the latter remains stronger on isolated shifts outside the four groups. The resulting fixed-weight policy transfers to hardware under payload changes, joint constraints, and leg extensions without online policy optimization. These findings identify recoverable expansion and domain composition as complementary considerations for learning generalizable motor control.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.