Which Tasks Train Adaptation? Paired Interventions for Meta-RL Task Selection
Abstract
Training a meta-RL agent requires tasks on which within-episode adaptation is useful, not merely difficult tasks. Existing curricula rank tasks by difficulty, regret, learning progress, or post-adaptation return, but these signals do not reveal how much performance depends on adaptation. To address this issue, we introduce Paired Adaptation-Contrast Ecology (PACE), which directly measures this dependence by comparing matched executions with the episode-local update enabled or disabled and evaluating both on a frozen suffix. The resulting learner-dependent return difference is used to update the training distribution in a closed loop as the learner changes. In controlled matched-Post experiments, high-AdaptGain task groups improve Post by 0.0069 more than matched low-AdaptGain groups under the same retraining budget. Across native benchmarks, PACE achieves the highest mean Post in 18 of 20 compatible benchmark–learner cells.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.