Evolving Agent Expertise from Conversational Feedback
Abstract
Deployed Large Language Model (LLM) agents accumulate extensive interaction logs in which user feedback implicitly signals where and how the model should improve. Learning from such feedback is difficult because the signal is both weak and mixed. It is weak because it arrives as free-form text with no success oracle, and because each remark covers only a fragment of what a task requires, so a reusable lesson emerges only after many remarks are aggregated. It is mixed because a single remark may point to a flawed procedure, a missing fact, or a skill applied where it does not belong. Existing methods each maintain a single kind of store and therefore absorb only part of this signal. Memory systems retain facts but never distill them into procedures, while skill-evolution methods distill procedures but discard the facts and never learn when a skill should not be applied. To address this, we present SkillWeave, which learns from feedback logs alone, without a success oracle or parameter updates, by jointly evolving a skill library, an evidence-grounded knowledge base, and the applicability conditions of each skill. Logs are processed in batches. Each batch first extracts missing facts, then groups logs into task topics, and finally creates or revises one skill per topic together with the conditions that govern its use. On user-feedback benchmarks covering legal, academic, and open-domain tasks, SkillWeave outperforms eight memory and skill-evolution baselines across three backbone LLMs, and it stays competitive on long-document tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.