NetHack+: A Configurable Environment for Long-Horizon Agent Research
Abstract
NetHack remains a challenging benchmark for long-horizon decision-making despite extensive research employing reinforcement learning, imitation learning, and language-model agents. A substantial difficulty in NetHack derives from the lack of infrastructure surrounding the game. As a result, we present NetHack+, an upgraded game engine with 42 knobs for varying game difficulty, 14× faster throughput than NetHack Learning Environment (NLE), and integrations into Gym, PufferLib, and Prime Intellect's Verifiers. We accomplish this by refactoring NLE to allow saving and restoring game state, configurable game mechanics, and custom observation and action interfaces embedded in modern LLM harnesses (Prime Agent and Claude Code). We also package a new metric, Bal-min, to reward balanced progression in dungeon depth and experience, motivated by successful human runs and designed to limit rewards for agents that descend without gaining experience. Experimentally, interface improvements and continual harness optimization yield modest gains in Bal-min, reaching approximately 1.5× baseline performance. To demonstrate the utility of checkpointing, we introduce Prime-Agent-Explore, an orchestrator-based search algorithm that selects saved checkpoints and guides subsequent exploration. Using 30 attempts with full vision, Prime-Agent-Explore reaches 11.38% Bal-min and 34.66% Bal-max, improving over the single-life baseline by 8.6× and 3.5×, respectively. With fog-of-war, Prime-Agent-Explore reaches 6.03% Bal-min and 30.93% Bal-max, exceeding the previously reported LLM Bal-min baseline by 3.3×. We also evaluate Prime-Agent-Explore on MiniHack, Crafter, and TextWorld, providing state-of-the-art performance across all three benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.