acceptodds
Under review as a conference paper at ICLR 2027

Model-Free Umbrella Reinforcement Learning

Abstract

We study hard reinforcement learning problems that combine sparse reward, state traps, and the absence of a terminal state. After an agent enters a trap, the trajectories it collects provide little evidence about an escape route. Umbrella Reinforcement Learning (URL) tackles this setting with an ensemble over the state space. Its objective combines return with the entropy of the ensemble's joint state and action distribution. The entropy encourages broad state coverage, while the return term favors rewarding behavior. The original URL formulation is written in continuous time and requires analytical state dynamics and their divergence. We introduce Model-Free Umbrella Reinforcement Learning (MF URL), a discrete-time instantiation trained from independent one-step queries to a resettable simulator, without analytical dynamics or policy trajectories, at the cost of resets to arbitrary states and many simulator queries. MF URL adapts the DualDICE objective by using a known proposal density as the denominator of the learned density ratio, so that the ratio determines the discounted state occupancy density. A second saddle point objective trains a critic with separate value and advantage heads. On Multi-Valley Mountain Car and StandUp, MF URL learns policies that reach and remain in the reward regions, with returns comparable to Model Based URL. Standard SAC and PPO remain well below both umbrella methods, even when PPO uses task-specific reward shaping. The learned densities recover the main structure of the Monte Carlo reference state distributions under the learned policies. A diagnostic comparison with MF URL SAC, which uses a temporal difference critic, finds that the saddle point critic agrees more closely with Monte Carlo reference returns under the policy learned by each method.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.