COME ON AGENT, LET’S STAY PLASTIC: WELL-CONDITIONED OPTIMISATION FOR MULTI- TASK AND META-RL
Abstract
Meta-learning with gradient adaptation assumes that pretraining on many tasks leaves a network that learns new tasks faster than a random one. We show that this can fail in deep reinforcement learning: networks pretrained on many tasks in parallel with Adam or Muon adapt to held-out tasks more slowly than a freshly initialised network. We ask whether optimisers that keep the weights close to orthogonal, which preserve plasticity when tasks arrive in sequence, also preserve adaptability after parallel pretraining. We study AdamO and MuonO, which add a decoupled soft orthogonality regulariser to Adam and Muon, and Parseval regularisation, and relate few-step adaptation to the neural tangent kernel of the pretrained network. In pre-registered experiments in two gridworld families, the regularised optimisers preserve the adaptability of deep networks and, at moderate width, adapt better than a fresh network, and they match or beat plasticity baselines and entropy controls. In the settings we tested outside gridworlds, pretraining did not reduce adaptability. The choice of pretraining optimiser therefore matters for gradient adaptation where pretraining reduces adaptability.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.