acceptodds
Under review as a conference paper at ICLR 2027

LOOPREPLAY: Reliability- and Diameter-Adaptive Sequence Replay for Rapid Recovery After Reward Relocation

Abstract

Under conditions where rewards are sparse and spatially distributed, three failures arise in sequence replay. Sweeping stalls due to iterative states, low-confidence stochastic edges contaminate propagation values, and a fixed replay horizon becomes mismatched with the scale of the learned graph. LOOPREPLAY introduces an event-based controller. This controller regulates access based on reward outcomes and signed prediction errors, enforces non-cyclic scans prioritized by empirical transition confidence, uses residual depletion as a termination condition, and determines propagation depth based on the diameter of the learned graph. All rules were fixed prior to running the validation experiment, which consisted of 12 conditions. This experiment covered three topologies (ring, bridge, cross), two sizes (9, 17), two slip probabilities (0.08, 0.22), and a balanced reward-shifting geometry. Compared with the stronger of prioritized replay and Dyna in each condition, LOOPREPLAY improved the average success AUC by +0.071 (95% hierarchical confidence interval [0.024, 0.126]) at matched replay compute per online step. The gain arises after the reward moves, where post-relocation success rose from 0.72 to 0.88. LOOPREPLAY remained within 0.05 AUC of the best baseline across all 12 conditions, and it also outperformed standard Dyna-Q, prioritized sweeping, and backward trajectory replay. The completion rate for backward sweeps exceeded 95% across all conditions, and the selection rate for low-confidence edges was less than 10%. The same rule improved all six conditions at the unseen size of 25. These results indicate that sequence reliability and the learned structural scale are joint control variables in replay-driven credit assignment.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.