Space-sampled Value Decay: Forgetting Mechanisms for Non-stationary Reinforcement Learning
Abstract
Reinforcement Learning agents deployed on physical systems must adapt continually, since degradation and shifting environment conditions change the dynamics (they drift) over time. In the hardest version of this problem, the agent interacts with a single system that might drift at every timestep, leaving no opportunity to revisit past conditions – a setting we call Single Environment, One-Shot Non-Stationary Reinforcement Learning (SEOS-NSRL). We argue that this setting calls for selective forgetting rather than re-learning, and introduce Space-sampled Value Decay (SsVD), which pulls value estimates of randomly chosen elements of the state space to a baseline value, so that outdated information in non visited regions is discarded. SsVD does not require resetting or change-point detection and plugs into modern off-policy algorithms; we integrate it into Soft Actor Critic and Deep Q-Networks. Across 6 non-stationary environments, SsVD improves upon its direct base algorithms and attains the best mean rank across all. The SsVD mechanism can also induce optimism which we show on hard-exploration tasks, although we investigate the connection here only briefly.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.