The Stored Action Does Not Identify the Sampled Latent: Finite-Precision Likelihood Replay in Squashed Policies
Abstract
A stored action is not always a faithful record of the sample that produced it. In a tanh-squashed policy, several large Gaussian samples can become the same float32 action, exactly or . Applying the inverse transform cannot recover the original sample. We study a Stable-Baselines3 PPO pathway that nevertheless reconstructs this value when evaluating action likelihoods. Its likelihood ratio initially equals one because collection and optimisation use the same reconstruction, but its gradient already differs from the gradient at the sampled value. Direct logging detects boundary collisions within the first eight environment interactions. Across three MuJoCo tasks, storing the sampled pre-tanh value substantially reduces peak approximate KL despite equal or greater boundary incidence. On HalfCheetah, the median falls from to . A separate PPO implementation reproduces the difference under a shared saturation stressor. We also show how a measured transform threshold predicts boundary exposure and how exploration initialisation changes that exposure. The practical distinction is between avoiding information loss and preserving the information needed for replay. Store the sampled latent when later updates evaluate its likelihood. When only a rounded action remains, model the probability of its rounding cell rather than treating a clamped inverse as the original sample.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.