acceptodds
Under review as a conference paper at ICLR 2027

Reward-Agnostic Replay for Shared-Structure Policy Evaluation

Abstract

How should stored transitions be prioritized when the same experience supports predictions of multiple signals? We study reward-agnostic replay for linear temporal-difference (TD) policy evaluation with a replay buffer. We exploit the decomposition of the expected TD update into a matrix shared across reward functions and a reward-dependent vector. By formulating estimation of the shared matrix as an approximate matrix multiplication problem, we derive a replay rule that prioritizes transitions using the product of the current-state feature norm and the norm of its discounted one-step feature difference. For the corresponding unbiased matrix estimator, this sampling rule minimizes expected squared Frobenius error within the estimator family we consider. The priorities are reward-independent, so they can be reused across multiple reward functions, including reward functions introduced later, without any reward-specific recomputation. In multi-reward evaluation with the reward-dependent vectors computed from the full buffer, a single norm-sampled matrix estimate improves value prediction over uniform replay and eventually outperforms reward-dependent PER as the number of tasks grows. The same priorities also transfer to held-out tasks and to tasks introduced later. Finally, we apply the same sampling rule to DQN on four MinAtar games, where we observe game-dependent gains in return and lower auxiliary multi-signal Bellman residuals in several settings. Together, these results help characterize when shared, reward-independent structure is useful for replay prioritization.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.