acceptodds
Under review as a conference paper at ICLR 2027

Experience Is Not Knowledge: Where the Decision Knowledge in an Agent's Memory Comes From

Abstract

LLM agents are commonly improved by writing a memory from their own trajectories and reading it back before the next episode. The gain is credited to experience, but the memory mixes what the writer already knew with what it induced from the trajectories. We ask where the useful decision knowledge in an agent’s memory comes from. The testbed is Slay the Spire, a roguelike deckbuilder: every run is generated fresh and the same decisions recur, so a rule written from a few runs is tested hundreds of times in fresh ones. We hold the player model, the combat executor, the prompt and the form of the memory fixed, and vary only the source of the memory’s content. The self memory is the player’s own summary of its games. The others hold outcome statistics from randomized simulator games, a stronger model’s prior knowledge alone, or, in the teacher memory, the player’s trajectories read by a stronger model under a few human-written guidelines. The player’s own memory does not raise its clear rate, and neither fewer games, a stronger writer nor writing one game at a time changes that. The memories written from statistics, from prior knowledge and by the teacher raise it, and the teacher memory raises it from 14 clears of 50 to 30. Replay at shared decision states shows that the player acts on its self-written memory, and that the memory poisons it. From the play counts in its own logs the writer concluded that the basic attack, which every deck starts with and so plays most, deserves an early upgrade, and the player then made in the exam an error it had almost never made in the games the memory was written from. The statistics and teacher memories reach their gains by different routes: one keeps adding cards to a large deck and the other stops. The teacher memory, written from one player’s trajectories, is the best memory for the other two players, whose own memories again do not help them. The three players converge on what it states simply and stay apart on the rules they must compose. Experience is not knowledge. The player improved when its memory held correct, usable rules, whether a stronger reader extracted them from the player’s own games or brought them from outside, and a wrong extraction from those same games became a new error.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.