acceptodds
Under review as a conference paper at ICLR 2027

Decoupled Policy Adaptation for Non-Stationary Reinforcement Learning

Abstract

Catastrophic forgetting is a key obstacle for reinforcement learning (RL) in non-stationary environments, where dynamics may repeatedly alternate across related regimes, causing adaptation to the current context to overwrite behaviors needed under previously encountered ones. Many prior methods rely on task identifiers or restrictive update constraints, making them less suitable for continuously varying contexts or online sample-efficient adaptation. We propose Decoupled Adaptation of Policies via HyperNetwork Interactions (DAPHNI), a sample-efficient architecture for continuous context-driven non-stationary RL that mitigates forgetting through decoupled, similarity-driven adaptation. DAPHNI decomposes a Gaussian policy into a state-only base policy to preserve shared behavior and a context-dependent residual modulator to localize context-specific adaptation. To handle continuously varying contexts, the residual modulator is implemented by a state-conditioned hypernetwork that induces multiplicative state–context interactions. A Kullback-Leibler (KL) divergence constraint further anchors the final policy to the base policy. We theoretically show that DAPHNI yields similarity-driven updates, transferring adaptation across aligned contexts while suppressing interference on dissimilar ones to mitigate forgetting. Comparison experiments and ablation studies show consistent gains over representative baselines across multiple benchmarks and diverse context dynamics, while mechanism analysis supports our design.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.