Skill-Forming Evidence for LLM Agents: Joint Selection and Target-Model-Aware Compilation
Abstract
LLM agents increasingly turn stored trajectories into reusable skills. Recent work has begun to show that trajectory composition can materially change downstream skill quality, alongside continued progress in skill extraction. We ask the next structural question: what makes a set of trajectories sufficient for forming a trans- ferable skill, and how should that evidence be represented for the target model? We formulate skill induction as two coupled decisions over experience: selec- tion, which chooses the trajectories that jointly identify transferable structure, and compilation, which expresses that selected evidence in a form the target model can use. In an exact Conditional Procedure World, six machine-verified proposi- tions show that skill-forming value can be irreducibly joint, that sparse support sets can preserve full-pool utility, and that the preferred evidence can change with the target model. Neural experiments reproduce these mechanisms across Pythia, Qwen2.5, and SmolLM2. Closed-loop tests then show that these data decisions matter for agent performance. On ALFWorld, changing the selected trajectories moves success from 0.164 to 0.351, and the gain from success-only selection replicates across five prospective experience pools. Raw delivery of the same failure-inclusive evidence reduces success by 0.157 relative to compi- lation. On TextWorld, a compiler that encodes episode-varying recipe content hurts transfer, whereas an invariant-aware compiler reverses the effect. Together, these results characterize skill-forming evidence as a jointly structured and target- model-dependent object: effective skills require informative trajectory sets and a representation that preserves their transferable structure for the model that uses the skill.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.