Learning to Distill Experience: Utility-Driven Experience Induction for LLM Agents
Abstract
Large language model (LLM) agents increasingly distill interaction trajectories into external memory to reuse past experience, yet determining what should actually be learned remains challenging. A single trajectory can support multiple plausible lessons, while its outcome provides limited evidence about which lessons will help in subsequent decision contexts. We formulate this challenge as utility-driven memory induction: learning to transform interaction trajectories into reusable knowledge by optimizing its downstream decision utility, rather than trajectory-level outcomes alone.. We propose EvoExp, a trainable experience inducer that generates actionable guidance paired with explicit applicability conditions. Following supervised initialization, Experience-Utility GRPO (EU-GRPO) optimizes the inducer through paired evaluations of a frozen actor under current and updated memory stores, using estimated action-value improvements across decision contexts as feedback. Once trained, EvoExp constructs new memories from fresh trajectories without further parameter updates or a deployment-time value evaluator. Across five strategic games, the RL-trained inducer outperforms the no-memory control in every game and achieves the best reported performance in three, including a +0.20 paired win-rate gain on Dou Dizhu and +10.30 chips on Texas Hold'em; despite RL training only on Dou Dizhu, EvoExp further improves on ALFWorld and GPT-authored Tide Ledger game.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.