acceptodds
Under review as a conference paper at ICLR 2027

GAMPO: Auditing and Optimizing LSTM Memory for Partially Observable Reinforcement Learning

Abstract

Recurrent policies must preserve useful evidence while limiting sensitivity to nuisance history. Forget-gate products omit paths through coupled LSTM cell–hidden dynamics. We use the complete joint-state Jacobian to bound finite history perturbations and define a full-path memory tail. Filter contraction and decoder calibration translate this tail into belief and control errors; a terminal-decision refinement avoids unnecessary horizon factors in recall tasks. Gate-Aware Memory Policy Optimization (GAMPO) combines an exponential teacher with sensitivity, information-state, and projection losses. Its population regret decomposition separates representation, critic, and projection errors, with explicit conditions linking fitted losses to these terms. Four controlled POMDPs reveal substantial limitations: the reported model-free variant remains near chance on recall tasks and can underperform unregularized sequence baselines. Full-path influence improves over a forget-only proxy on active sensing, and an independent fast-refresh control shows a regret benefit over omission. Short-delay and exact-belief controls expose persistent learning limitations even when memory or representation demands are reduced. These findings support full-path sensitivity as an auditing tool and expose the additional requirements for effective memory regularization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.