Understanding Information Freshness in Reinforcement Learning under Diffusive Distribution Shift
Abstract
Reinforcement learning relies on past interactions with the environment, but distribution shift can change how useful historical data remain for current decisions. Existing nonstationary-RL models often measure the cumulative magnitude of environmental change. We introduce diffusive distribution shift, a natural model for many real-world systems whose dynamics evolve continually through stochastic changes that may accumulate, cancel, or revert over time. We establish a simple, quantitative information freshness principle: historical data should be retained until the decision error caused by changing dynamics becomes comparable to the statistical uncertainty reduced by keeping them. This principle yields explicit expressions for how long historical data remain useful and for the unavoidable cost of tracking changing dynamics. Guided by this principle, we develop optimistic model-based RL algorithms for both known and unknown diffusion rates. We prove matching upper and lower bounds, up to logarithmic factors, for the additional regret caused by distribution shift, establishing the minimax optimality of freshness-guided learning. Moreover, without knowing the diffusion rate in advance, our adaptive algorithm achieves the same regret rate, up to logarithmic factors, as the algorithm tuned with the true rate. Controlled experiments validate the predicted freshness mechanism, and a wind-storage study illustrates its decision-level consequences.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.