SkillTrainBench: Can LLM Agents Author Skills That Improve Other LLM Agents?
Abstract
Skills are compact packages of instructions and reference material supplied to an LLM agent's harness, allowing its behaviour to change without updating model weights. We ask whether LLM agents can write such skills for other LLM agents. SkillTrainBench gives a curator agent a fixed interaction budget with one or more learner agents whose model weights remain fixed. The curator iteratively writes and revises a skill, chooses which learner-task pairs to probe, and uses the resulting performance to guide further revisions. The final skill is evaluated on private held-out tasks. We study four domains drawn from HealthBench, Humanity's Last Exam, quantitative-finance evaluations, and τ³ Banking, using multiple frontier curators and learners spanning a range of capabilities. We find that curator-written skills improve every learner on HealthBench and transfer to models that never saw them, but rarely help on the other three domains. Skill quality is also non-monotonic: performance frequently peaks before the final iteration, and curators rarely submit their best skill because the development scores they see cannot rank their own drafts. SkillTrainBench therefore provides a controlled setting for studying how LLM agents can improve other LLM agents through externally authored skills, and where this process currently fails.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.