acceptodds
Under review as a conference paper at ICLR 2027

TopoExplore: Homology as an Exploration Signal

Abstract

Exploration signals in reinforcement learning are currently computed from what an agent has seen: visitation counts, density estimates, or a model's prediction error at individual states. None of these report the holes in the visitation space. To capture them, we compute the persistent homology of the agent's own archive of visited states during exploration and turn each detected class—an enclosed region the archive surrounds but has not entered—into a selection bonus concentrated on the archived states from which entry is possible, gated by an attempt counter that retires candidate entrances and prunes sealed structures. The method, TopoExplore, is Go-Explore plus one additive term, so the comparison with Go-Explore isolates that term. On the held-out worlds of the open-source TopoGym benchmark of topologically varied hard-exploration environments, TopoExplore finds the goal in worlds against for Go-Explore and, where both find it, with a median k steps against k. On an environment that stress-tests far-apart chambers that create holes in the archive space, TopoExplore enters all six or eight chambers across seeds, against and of seeds for Go-Explore. On Montezuma's Revenge, built over the room graph, the term helps whenever the archive surrounds a room it has not entered, puts –% of the selections on the surrounding rooms while it is active, and enters the room sooner in five of six cases; such enclosures are rare on this game, so the score is unchanged. Beyond discrete, restorable state spaces, the same signal drives an embodied explorer that walks to every target it selects and explores from depth alone; a vision-language-action navigation policy fine-tuned on its trajectories succeeds % more often on unseen scenes than one fine-tuned on shortest-path oracle demonstrations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.