Ask the Memory, Not the Teacher: Episode-Swap Gating Prevents Stale-Memory Hallucination in On-Policy Distillation of Agent Memory
Abstract
Agents that keep a memory of past episodes are increasingly consolidated into their weights. A teacher that reads a retrieved episode supervises, token by token, a student that does not, until the student no longer needs the store. We show that this on-policy consolidation writes more into the weights than it should. A retrieved episode mixes the procedure that solves a family of tasks with values that held only in that episode, such as an access code or the place where an object happened to be, and dense distillation transfers both. The consolidated student then states a past episode’s credential as a fact about a new user, with no retrieval record left to trace it. Standard token-selection signals miss this failure, because a teacher that copies from its context is confident, and the copied values can appear anywhere in a response. We propose to test each token against the memory instead of trusting the teacher. When the retrieved episode is replaced by a sibling episode of the same task family, the teacher’s predictions of procedural tokens stay the same while its predictions of episode-bound tokens change, and we use this difference as a per-token distillation gate that costs two extra teacher forward passes with one sibling. In a diagnostic environment where the provenance of every token is known, the gate separates the two kinds of tokens almost perfectly while confidence and position heuristics barely beat chance. During consolidation it eliminates stale-value emission at equal task success, nearly doubles honest abstention, and carries over to a second model family and to embodied agents in ALFWorld.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.