acceptodds
Under review as a conference paper at ICLR 2027

Balancing Learnability and Bellman-Error Coverage in Prioritized Experience Replay

Abstract

Experience replay improves the sample efficiency of off-policy reinforcement learning, but its benefit depends on which transitions are revisited. Prioritized experience replay (PER) favours large temporal difference (TD) errors, which may overemphasize noisy or irreducible samples. Reducible-loss prioritization (ReLo) instead targets errors that appear learnable, yet it can fail in the complementary direction: when the reference loss is imperfect, difficult but informative transitions may receive negligible priority and disappear from replay. We introduce BLCR (balanced learnability–coverage replay), a conservative repair that combines a normalized TD-error priority floor with bounded reducible-loss weighting. Unlike an additive mixture, the floor is activated only when ReLo underprioritizes a high-error transition; otherwise, the reducible-loss ranking is preserved. Across five MinAtar games at two million frames and five seeds, BLCR improves over ReLo on every game and obtains the highest final return on Asterix and Breakout. Across 30 ALE-RAM games, it outperforms ReLo on 25 games and wins 118 of 150 paired-seed comparisons. In a preliminary 300K-step SAC extension over four MuJoCo locomotion tasks, BLCR-SAC ranks first on Hopper and Walker2d, improving the final return over the SAC by 12.8% and 56.5%, respectively, while the results on Ant and HalfCheetah remain mixed. These findings position conservative TD error coverage as a targeted correction to reduce loss replay rather than a universal replacement for PER.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.