Self-Generated Feedback Destabilizes Test-Time Training: A Causal Decomposition of Long-Horizon Adaptation
Abstract
Test-time training (TTT) gives a model writable state: information can persist in its weights after the source text leaves the attention window. The same mechanism creates feedback when the model learns from its own output: each update changes the model that generates the next training example. We study this feedback over 128K-token streams. Retaining generated-text writes worsens prediction of independent human-written text at three TTT-E2E scales (125M, 760M, and 3B), and the same failure occurs when Adam updates Qwen3-4B’s existing weights. Three matched comparisons provide a causal decomposition of the result. First, Fixed Generation, in which a frozen model generates every training chunk, removes 97.5% of the measured damage even though the adapting model continues to update. Second, Recorded Replay of the same degraded text shows that reading it through attention causes one loss, while retaining its updates stores an additional loss in the weights. Third, branching from an identical state before one update reveals the local conflict: the update predicts its source chunk better but predicts new real text worse. This cost grows after closed-loop adaptation, and a few source trajectories account for most large failures. Settlement evaluates the complete candidate state on independent real text before commitment, reducing the endpoint gap to at most .07 nats at 125M and 760M while retaining real-text adaptation. Thus, optimizing the current self-generated example can make later, independent inputs harder to predict.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.