First-Explore+: Sample-Efficient Training of In-Context Exploration
Abstract
Exploration is essential for learning, yet common exploration strategies—such as random action selection, generic exploration heuristics, or priors acquired through imitation learning—adapt slowly to prior attempts, can be highly domain specific, and often require orders of magnitude more interactions than a human to systematically explore the same space. In-context meta-reinforcement learning can learn to explore adaptively and as efficiently as a human, but can also fail entirely in environments with deceptive rewards. First-Explore addresses this by meta-learning separate exploration and exploitation policies, but its original instantiation is highly sample-inefficient, restricting experiments to simple environments and potentially obscuring phenomena that only emerge at greater scales. We introduce First-Explore+, which substantially improves the efficiency of First-Explore by adding a value baseline, incorporating better transformer and RNN architectures, and developing several extensions to the First-Explore framework. First-Explore+ matches or exceeds the original method's performance in challenging meta-learning benchmarks with deceptive rewards using orders of magnitude fewer training samples. This increased efficiency allows us to evaluate the First-Explore framework in a substantially more complex environment, XLand Mini-Grid, where First-Explore+ achieves high scores while other meta-learning baselines completely fail to explore. In this environment, we provide the first clear evidence of a new phenomenon—explicitly meta-learning cross-episode exploration can enable successful zero-context single-episode performance where an otherwise matched single-episode baseline fails entirely. These results provide preliminary support for a broader role of meta-RL: training a policy for cross-episode exploration to improve returns in standard single-episode RL. Together, these advances represent progress toward increasingly general exploration priors and potentially a key step toward a foundation model for general exploration.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.