Do Not Disturb: When Plasticity Interventions Degrade Offline-to-Online Reinforcement Learning
Abstract
Maintaining neural network plasticity is widely regarded as an indispensable pre- requisite for continual reinforcement learning (RL). A vibrant literature advocates periodic parameter interventions—such as network resets, Shrink-and-Perturb, and dormant neuron recycling—to counteract capacity loss, dead units, and feature rank collapse. While effective in continual supervised learning and long-horizon online RL from scratch, these methods rest on an implicit assumption: that plasticity restoration is universally benign and can be scheduled unconditionally. In this work, we demonstrate that this assumption fails fundamentally in offline-to-online (O2O) continuous-control RL. When transitioning from static pretraining to online interaction, agents inherit structured, high-performing policy and value manifolds. Across standard D4RL continuous-control locomotion benchmarks, we show that unconditional peri- odic Shrink-and-Perturb induces severe intervention vulnerability, destabilizing converged locomotion policies and causing a catastrophic 68.3% collapse in ag- gregate Interquartile Mean (IQM) normalized return (38.78 → 12.30, paired Wilcoxon signed-rank W = 18.0, p = 1.44 × 10−11, mean paired difference +26.03 ± 19.09). To address this vulnerability, we articulate the “Do Not Dis- turb” principle: an agent should intervene only when representation capacity is demonstrably impaired. We introduce CAPACITYGATE, a closed-loop diagnostic framework that monitors feature effective rank and neuron dormancy using uni- form replay reservoir sampling and refractory cooldown control. In stable and moderately shifted transfer, CAPACITYGATE detects that representation capacity remains intact (ρ ≈ 1.0, d < 0.08) and safely abstains from intervention, preserv- ing baseline performance (38.78 IQM) and protecting converged policies against perturbation-induced collapse. Furthermore, under non-stationary physical stress (actuator crippling), we uncover a profound operator stability asymmetry: zero- functional-shift neuron recycling (ReDo) preserves bipedal balance manifolds (8.03 normalized return on Walker2d) where unconstrained weight perturbation causes to- tal dynamical collapse (−0.53). Our findings reframe plasticity restoration in O2O RL from an open-loop maintenance routine to a regime- and operator-dependent intervention that requires rigorous diagnostic gating.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.