PINNs Meet Umbrella Reinforcement Learning: Improved Convergence and Training Consistency
Abstract
Hard reinforcement learning problems combine sparse or delayed rewards, state traps, and the absence of a terminal state. Umbrella Reinforcement Learning (URL) addresses them with a continuous ensemble of agents described by a discounted state occupancy density and a steady-state value function, which enter the policy update through an entropy-augmented reward and a modified advantage. The original URL implementation trains these two quantities with stochastic semi-gradient updates that do not directly enforce their governing equations. We introduce PINN-URL, a physics-informed training scheme that learns both of them by minimizing their squared strong-form residuals, using sinusoidal representation networks (SIRENs) and exact sums over discrete actions, while retaining the URL policy update. On Multi-Valley Mountain Car and StandUp, PINN-URL converges faster and to a higher final return, with substantially lower run-to-run variability. Tuned PPO, including long-horizon, reward-shaped, and locally imitation-assisted variants, does not reach comparable return at the target discretization within its interaction budget. In representative runs, PINN-URL's learned value function and occupancy density agree more closely with policy-dependent simulation references. Ablations show that the joint state–action entropy is essential on Multi-Valley Mountain Car, and that neither the network architecture nor the learning configuration alone is advantageous on both benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.