acceptodds
Under review as a conference paper at ICLR 2027

The Reset-to-Rollout Gap in Adaptive Sampling for Humanoid Motion Tracking

Abstract

Recent advances in humanoid motion tracking have enabled policies to imitate diverse human motions with high fidelity. Learning such repertoires requires allocating training samples across motions of varying difficulty. Probability-based adaptive sampling blends failure-based reset prioritization with uniform sampling. Weaker blending favors faster improvement in overall tracking quality, whereas stronger blending promotes difficult-motion acquisition but can still leave some motions slow to learn or unlearned. We find a reset-to-rollout gap: difficult regions can contribute a substantially smaller share of rollout samples than their reset probabilities suggest. Increasing reset priority alone therefore does not ensure sufficient local practice for difficult-motion acquisition. To address this gap while combining weaker blending's faster refinement with stronger blending's acquisition benefits, we introduce Adaptive Rehearsal Allocation (ARA). This plug-in reserves a bounded pool of parallel environments for repeated rehearsal of difficult regions, while the remaining environments retain ordinary adaptive sampling. Across the evaluated settings, ARA consistently acquires all target motions. In training from scratch, weaker adaptive sampling augmented with ARA acquires all target motions and achieves better tracking quality than stronger adaptive sampling alone under the same training budget.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.