acceptodds
Under review as a conference paper at ICLR 2027

ASCENT: Online Test-Time Training of Long-Horizon Agents via Self-Distillation of Verified Experience

Abstract

A large language model (LLM) agent solves long-horizon tasks through many reasoning-action turns, with one verification signal at termination. Deployed agents face streams of related tasks, making their own trajectories a natural resource for improvement. In-context evolving agents store reflections, memories, or skills as text, so reuse depends on retrieving the right experience and on a frozen policy executing it. We study Online Agentic Test-Time Training (OaTTT), which trains the LLM during deployment on its own execution trajectories. Each task is attempted once, and the executed trajectory with its verification result is the only learning data for updates that persist to later tasks. Directly imitating the sampled tokens of this single rollout destabilizes the policy. We introduce ASCENT (Agentic Self-distillation for Cross-task EvolutioN at Test-time), which instead self-distills verified experience. A frozen copy of the LLM receives the verifier-accepted trajectory, with environment-rejected turns removed, as privileged information and predicts next-token distributions along it with this hindsight; distilling them into persistent LoRA fast weights updates the agent for later tasks, without an external reference solution or stronger teacher. We characterize the population target of this distillation and the limits of sparse outcome selection. Across ALFWorld, WebShop, and AppWorld on varied model scales, ASCENT improves task success and interaction efficiency as experience accumulates, outperforms the compared online adaptation methods, and transfers to held-out scenes, showing that a deployed agent can consolidate its own verified experience into its policy without a separate training phase or memory retrieval at inference.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.