The Reset-to-Rollout Gap in Adaptive Sampling for Humanoid Motion Tracking
Abstract
Recent advances in humanoid motion tracking have enabled policies to imitate diverse human motions with high fidelity. Learning such repertoires requires allocating training samples across motions of varying difficulty. Probability-based adaptive sampling blends failure-based reset prioritization with uniform sampling. Weaker blending favors faster improvement in overall tracking quality, whereas stronger blending promotes difficult-motion acquisition but can still leave some motions slow to learn or unlearned. We find a reset-to-rollout gap: difficult regions can contribute a substantially smaller share of rollout samples than their reset probabilities suggest. Increasing reset priority alone therefore does not ensure sufficient local practice for difficult-motion acquisition. To address this gap while combining weaker blending's faster refinement with stronger blending's acquisition benefits, we introduce Adaptive Rehearsal Allocation (ARA). This plug-in reserves a bounded pool of parallel environments for repeated rehearsal of difficult regions, while the remaining environments retain ordinary adaptive sampling. Across the evaluated settings, ARA consistently acquires all target motions. In training from scratch, weaker adaptive sampling augmented with ARA acquires all target motions and achieves better tracking quality than stronger adaptive sampling alone under the same training budget.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.