Retracing the -Discrepancy: Memory Learning from Off-Policy Data
Abstract
Acting optimally under partial observability requires memory. Allen et al. (2024) recently proposed to use the -discrepancy, defined as the gap between value estimates obtained at different bootstrapping horizons, as a model-free signal for learning memories that disambiguate latent states. However, the on-policy formulation limits its applicability when fresh interaction data is expensive. Extending the -discrepancy to off-policy learning is nontrivial. When trajectories are generated by a different policy, the resulting multi-step returns no longer correspond to those of the target policy. Importance sampling can correct this mismatch, but often at the cost of prohibitively high variance. We overcome these issues by introducing the Retrace discrepancy, an off-policy generalisation of the -discrepancy built from Retrace-style trace corrections (Munos et al., 2016). We first show that the corresponding Retrace operator admits a unique fixed point even without assuming that the observations are Markovian. We then prove that the resulting discrepancy preserves the key partial-observability detection property of the original -discrepancy: if a POMDP has a nonzero on-policy -discrepancy for some policy, then, for almost every combination of behaviour and target policy in the off-policy setting, the Retrace discrepancy is nonzero as well. Experiments demonstrate partial-observability detection and memory learning from replay. On Battleship and Minesweeper, the replay-trained agent first exceeds selected mean-return thresholds with 4.4x and 3.0x fewer environment interactions, respectively, than PPO with the on-policy -discrepancy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.