AIPR:Adaptive Isometric Plasticity Regularization for Reinforcement Learning
Abstract
Deep reinforcement learning encounters a balance breakdown between stability and plasticity during long-term training, which causes a degradation in the network's continuous optimization capability and impairs the agent's long-term learning ability. Some existing approaches employ the Deviation from Isometry (DfI) to measure how far a weight matrix strays from an isometric structure. However, discrete parameter reinitialization interrupts the continuous optimization process, while abrupt parameter updates may disrupt the policy distribution. To address this issue, we propose Adaptive Isometric Plasticity Regularization (AIPR). This method transforms DfI from a trigger metric for discrete parameter reinitialization into a continuously differentiable geometric regularization term. By jointly optimizing this term alongside the reinforcement learning task objective, AIPR continuously constrains the evolving geometric state of network weights throughout training. Concurrently, AIPR incorporates a DfI-aware adaptive gating mechanism that dynamically regulates regularization intensity based on the network's current weight geometric state. Considering potential discrepancies in weight geometry between the Actor and Critic, AIPR separately characterizes their geometric states, generates corresponding adaptive regularization weights, and preserves the original task coupling within the Actor-Critic framework. AIPR is compatible with both on-policy and off-policy learning, enabling seamless integration with PPO and SAC. Empirical results demonstrate that AIPR achieves competitive learning performance across multiple continuous control benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.