Resolving Non-Convergence in Distributional Control: A Homotopy Approach with Double Distributional Value Iteration
Abstract
Distributional reinforcement learning models the full probability distribution of cumulative returns rather than just their expected values. In control settings, standard algorithms select actions using a greedy policy based on the mean returns of the current distributional estimates. However, when multiple actions have the same optimal expected return but different return distributions, small estimation errors in their means cause the greedy policy to abruptly jump back and forth between these actions. Because the policy keeps switching, the target return distributions for Bellman updates oscillate continuously, preventing the estimated return distributions from converging even after value errors become negligible. To resolve this chattering issue, we provide an algorithm,at each step, an auxiliary update estimates action values and turns them into a softmax reference policy, and a main update learns the return distribution by distributional Bellman evaluation under that policy. The softmax is evaluated at a positive temperature λt. Cooling the temperature to zero throughout training brings the policy back to the optimal actions and, on a tie, shares its mass uniformly over them, while the objective and the return target stay unchanged. Theoretically, for tabular MDPs, we prove for the planning version of our algorithm that the softmax policy converges to the uniform distribution over optimal actions and that the return distributions converge to its law in maximal 1-Wasserstein distance. For its stochastic approximation version, we prove almost-sure convergence to the same policy and return law. We provide a practical implementation using categorical networks. Experiments in tabular environments confirm that algorithm eliminates distributional oscillation and guarantees convergence, while continuous-control benchmarks in MuJoCo demonstrate improved distributional stability during training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.