From Hesitation to Intuition: Internalizing Test-Time Adaptation into LLM Agents via On-Policy Self-Distillation
Abstract
Memory-augmented large language model (LLM) agents adapt at test time by retrieving past experience into their context, thereby avoiding the prohibitive cost of updating parameters online. Existing approaches, however, often suffer from two fundamental limitations: the decision of when to refresh the strategy relies on the agent's own self-evaluation, which is vulnerable to cognitive hallucinations, and the acquired experience remains entirely outside the model parameters, leading to repeated memory retrieval and strategy reconstruction. To address these issues, we propose TMA, a Thinking Mode-Aware agentic memory framework that consolidates test-time adaptation into the policy parameters. Specifically, we learn a linear discriminant probe on the hidden states of confident and hesitant chain-of-thought (CoT) segments, where hesitant segments exhibit typical failure patterns, and project the agent's hidden state onto the learned direction as a hesitation score. Reflection is triggered only at hesitant states, refining the global strategy with retrieved trajectory memory. We further develop an on-policy self-distillation scheme, where a teacher conditioned on privileged hindsight, i.e., successful trajectories and high-quality reflections, supervises the student on its own rollouts to jointly optimize task execution and reflection generation. This enables slow reflection to be progressively absorbed into fast single-pass decisions, reducing subsequent reflection triggering. Experiments on ALFWorld, WebShop, and Jericho demonstrate that TMA outperforms state-of-the-art memory baselines with a relative success-rate gain of up to 8.7%. After five rounds of self-distillation on ALFWorld, TMA further improves the success rate by 6.4%, while reducing reflections and token consumption per task by 92% and 28%, respectively.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.