Vivid Memories: Fine-Grained Trajectory Memory for On-Policy Agent Self-Distillation
Abstract
Memory is central to self-improving language agents, but most agent memory systems compress experience into skills, procedures, or reflections. This is useful for inference-time prompting, but it discards the step-level state–action evidence needed for training-time credit assignment. We introduce VIVID, an on-policy self-improvement method that treats memory as a training-time supervision bank rather than an inference-time prompt aid. VIVID stores both successful and failed rollouts as raw turn sequences, annotates each turn with an offline outcome label and a continuous value, retrieves a balanced mixture of state-matched successful and failed precedents during training rollouts, and uses the annotations for reward-consistent turn-level advantage shaping of GRPO: the direction of every update is inherited from the verified environment reward, and the judge modulates only its magnitude. Evaluation is retrieval-free: the trained policy uses no memory, retrieval, or judge. On AppWorld, ALFWorld, and Sokoban with Qwen3-1.7B, Qwen3-4B, and Llama-3.1-8B backbones, VIVID is best on every benchmark, metric, and backbone, improving Qwen3-4B average accuracy from 52.5 (GRPO) and 64.1 (Skill-SD) to 72.0 (3 seeds, std 1.0). Controlled analyses locate the gain in the content of per-turn labels rather than in reweighting machinery or external-LLM spend: uninformative-label controls collapse to the GRPO floor, judge-free retrieval already matches Skill-SD with zero external LLM, budget-matched baselines do not close the gap, a future-blind (causal) judge retains most of the improvement, and letting the judge set the update sign is measurably worse than magnitude-only shaping.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.