BEFORE THE COLLAPSE: MEASURING AND RESTORING PLASTICITY IN CONTINUAL VISION-LANGUAGE-ACTION REINFORCEMENT LEARNING
Abstract
Sequential RL post-training of large vision-language-action (VLA) policies has recently been reported to be surprisingly stable: with LoRA adapters and on-policy GRPO, near-zero backward transfer is achieved across standard benchmarks, and stability is attributed to scale, low-rank constraints, and pretraining. We argue that this focus on forgetting inverts the field's priorities: low forgetting can coexist with, and even mask, insufficient plasticity (the capacity to keep learning), yet current evaluations measure only behavioral success, never network-level trainability. We present a closed loop of measure, intervene, predict for continual VLA RL. On a five-task stream of increasing difficulty, a standard forgetting-oriented anchoring objective suffers a catastrophic policy collapse at the fourth stage: success on all ten tasks falls to zero simultaneously, preceded by a structural fingerprint (effective-rank contraction, unbounded norm growth) that is absent in the unanchored baseline, whose explored subspace instead expands. The fingerprint prescribes the remedy: DARE-style sparsification of the collapsed adapter restores acquisition at the collapse stage () with retention within pp of the baseline, and directional rank-channel release is therapeutic ( random release) while improving the healthy frontier; both replicate across seeds. A lightweight forgetting-risk predictor built on representation similarity and NTK-overlap theory rank-orders streams by their actual forgetting behavior. These deliverables turn the reliability of continual VLA RL from an emergent accident into an engineered property.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.