Self-Induced Distribution Shift in Agent Memory: Write-Time Retention Criteria Cannot Be Optimal
Abstract
Every published retention criterion scores a memory item by an estimate of how much it will matter later, computed under a different memory policy from the one the criterion is about to install. This is not a sampling nuisance to be reduced with more rollouts: the channel that compresses decides which screens the agent reaches, hence which facts are later needed, hence what it should have kept — the criterion's own output moves the distribution its score was computed against. We formalise the induced-distribution operator, prove its fixed point exists by Brouwer and is unique under a contraction condition, and show that the objective every write-time criterion actually optimises is exactly the first iterate of that operator, with its error governed two-sidedly by one measurable quantity: the shift between the law the criterion assumed and the law its own output induces. Optimality then requires that shift to vanish, which requires memory not to affect behaviour. We claim nothing for the operator: it is performative prediction's decoupled-risk map and the single-agent endogenous setting is performative reinforcement learning's. What is ours is the feasible set — a rate-constrained code, whose argmin jumps where a parametrised policy's moves smoothly, so solution stability is a hypothesis here and a mild condition there — and a quantitative failure needing neither: a two-point construction on which every write-time criterion, deterministic or randomised, pays at least half of sigma in performative risk at retention strength sigma, both bounds attained. Results go against us twice. The fixed point these criteria fail to reach is itself stable, not optimal, leaving 24.0% of the performative risk unclaimed — which raises the wall rather than lowering it. And we withdraw a theorem of our own: the separation we claimed grows with the horizon is a window, exactly zero at both ends. On 19 real application flows walked on devices, a 7B agent's law shift reaches 3.28x its measured noise floor at a quarter budget, moving 17 of 19 flows, with the mechanism visible: the revisit rate jumps from 0.38 to 0.497 because a code that forgets what was already done sends the agent back to do it again.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.