Snapshot Entropy Exploration in Reinforcement Learning
Abstract
Maximum state-entropy exploration seeks broad state visitation, but the standard visitation distribution pools variation across both time and independent executions. Consequently, the same pooled coverage can arise either from a reproducible, time-structured route or from dispersion across runs. We make this distinction explicit by sampling a time according to the visitation weights, independently of a run, and letting denote the agent's state at that time. Then is the pooled state entropy, while the chain rule separates it into two components. We call the snapshot entropy: it measures the uncertainty across executions that remains once elapsed time is known. We study exploration with certainty, which maximizes pooled state entropy subject to a snapshot-entropy budget. This budget limits time-aligned disagreement across runs and guarantees that expected within-run visitation entropy is at least the pooled state entropy minus the budget. We derive an exact policy-gradient identity for this objective and obtain Certainty Policy Gradient (Certainty PG), a model-free method that estimates time-indexed state distributions from synchronized rollouts. We characterize the structure and limits of exploration with certainty. Deterministic policies suffice to minimize snapshot entropy; selecting a high-coverage certain route is NP-hard; and stochastic transitions impose an unavoidable uncertainty floor computable by dynamic programming. On Ant, a continuous-control benchmark used in prior maximum-entropy exploration work, synchronized rollouts expose same-time dispersion hidden by pooled occupancy: Certainty PG reduces this dispersion, though at a cost in coverage. On clean Rooms-8, where broad reproducible routes exist, Certainty PG retains 98.0% of Pooled PG's pooled entropy while increasing single-run all-state coverage from 31.5% to 99.8%. A trajectory-entropy objective achieves similarly high coverage but substantially greater time-aligned uncertainty, demonstrating that single-run coverage and repeatability are distinct objectives. On noisy Rooms-8, it approaches the uncertainty floor, while navigation stress tests show that excessive concentration can produce incomplete routes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.