Reward-Agnostic Replay for Shared-Structure Policy Evaluation
Abstract
How should stored transitions be prioritized when the same experience supports predictions of multiple signals? We study reward-agnostic replay for linear temporal-difference (TD) policy evaluation with a replay buffer. We exploit the decomposition of the expected TD update into a matrix shared across reward functions and a reward-dependent vector. By formulating estimation of the shared matrix as an approximate matrix multiplication problem, we derive a replay rule that prioritizes transitions using the product of the current-state feature norm and the norm of its discounted one-step feature difference. For the corresponding unbiased matrix estimator, this sampling rule minimizes expected squared Frobenius error within the estimator family we consider. The priorities are reward-independent, so they can be reused across multiple reward functions, including reward functions introduced later, without any reward-specific recomputation. In multi-reward evaluation with the reward-dependent vectors computed from the full buffer, a single norm-sampled matrix estimate improves value prediction over uniform replay and eventually outperforms reward-dependent PER as the number of tasks grows. The same priorities also transfer to held-out tasks and to tasks introduced later. Finally, we apply the same sampling rule to DQN on four MinAtar games, where we observe game-dependent gains in return and lower auxiliary multi-signal Bellman residuals in several settings. Together, these results help characterize when shared, reward-independent structure is useful for replay prioritization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.