acceptodds
Under review as a conference paper at ICLR 2027

MineMem: A Benchmark for Experience Transfer and Updating in Interactive 3D Worlds

Abstract

Vision-language model (VLM) agents can interpret visual observations and follow language instructions, but long-term tasks also require them to reuse past experience and revise it when the environment changes. Existing memory benchmarks often measure task success without revealing whether an agent applies reusable experience or merely recalls facts from a demonstration. To bridge this gap, we introduce MineMem, a benchmark built on the open-source voxel engine Luanti for evaluating how VLM agents transfer and update experience in interactive 3D worlds. Agents acquire experience from human demonstrations and then perform autonomous tasks under controlled environment transformations. MineMem evaluates five capabilities: spatial reasoning, object-state tracking, temporal event reasoning, planning, and memory updating. Its procedurally extensible templates vary demonstration length, event count, action-dependency depth, spatial displacement, object-state changes, and conflicting evidence, enabling controlled tests of memory horizon and transfer difficulty. We compare complete memory mechanisms with no-history and fact-injection controls, measuring task success, verified progress, and resource use. Contrasts across capabilities and environment transformations are designed to reveal when retained experience supports transfer and when stale or conflicting experience causes negative transfer. The benchmark code will be released upon acceptance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.