acceptodds
Under review as a conference paper at ICLR 2027

Replay-Conditioned Explanations: Disentangling Current, Historical, and Constraint Effects in Memory-Based Reinforcement Learning

Abstract

Explaining recorded decisions of memory-based reinforcement-learning policies requires distinguishing current observations, retained history, and action constraints. We introduce replay-conditioned explanations, which bind each target decision to its policy and replay context and separately evaluate current-input replacement, historical intervention with suffix replay, and changes to action legality. The contributions of this paper are an executable query contract and an empirical characterization of the gap between computational validity and reference-dependent fault diagnosis. Reconstruction checks and optional exact Shapley attribution make the evaluated computation explicit. We validate query correctness on analytical mechanisms, reward-trained recurrent policies with 8- and 64-step horizons, and a native POPGym task with a three-step memory dependency. Public TimeSHAP agrees with exact event-level attribution under matched queries. In the independent-reference study, recovery falls from 100% with clean paired references to approximately 52% under the tested retrieval rule. A gate calibrated on clean episodes reduces clean-context modifications from approximately 48% to 4–5%, while recovery falls to 16–17%. Greedy search matches exhaustive recovery with fewer policy evaluations on these tasks. These results support explicit query specification and reference checks, while showing that computational correctness alone does not establish diagnostic reliability. Broader benchmark generalization and new industrial validation remain open.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.