SkillMuse: Multi-Turn User Simulation as a Feedback Generator for Skill Evolution
Abstract
In enterprise conversational settings such as customer support, LLM-based assistants rely on high-quality domain skills to answer; yet no mechanism attributes execution failures to skill defects and turns them into targeted improvements—operational data keeps accumulating while skill quality stagnates. Recent work on skill self-evolution builds a revision loop toward closing this gap, but the feedback driving the loop is derived from single-turn question-answering evaluation: a single exchange exposes only the gaps in the user's opening statement, so once the first revision has patched them, the evaluation supplies little further signal, and defects that surface only across multiple turns remain unreached—the loop is intact, yet evolution stalls. Governance in these systems likewise rests on a scalar score, which can only reject a degraded candidate outright, without diagnosing or repairing its structural cause. We argue that the binding constraint on sustained skill evolution lies less in editing capability or the number of iterations than in whether the evaluation feedback keeps supplying trustworthy evolution gradients. We introduce SkillMuse, coupling two components. Trustworthy feedback generates the gradient: we recast multi-turn user simulation from an evaluation endpoint into a feedback generator, so follow-up questions expose defects layer by layer and every round of revision both consumes and produces feedback. Controllable governance constrains the gradient direction: an independent governance layer actively repairs factual degradation and structural bloat after each revision, so degradation does not accumulate across rounds. Across six categories of cloud services, nine production skills, and 98 skill-reference files, SkillMuse exceeds self-reflection-based evolution by 23.0 points and single-turn-QA-driven evolution by 15.4 points.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.