Forgetting is What to learn: Verifier-Grounded Minimal-Core Skills for Self-Evolving Agents
Abstract
AI for Science (AI4Science) increasingly relies on AI agents to execute long-horizon scientific workflows that require multiple reasoning steps, specialized tools, data transformations, and domain-specific procedural constraints. Reusable skills can transfer useful scientific experience across tasks, but directly accumulating successful workflows also preserves redundant, source-specific, and behaviorally unnecessary procedures. We present AutoSkills, a verifier-grounded framework for evolving reusable skills for long-horizon task execution. Given a successful reference workflow, we construct an over-complete skill and performs verifier-grounded forgetting: candidate operations are removed and the solver is rerun under the same executable verifier, retaining only deletions that preserve task performance. The remained core skill is then generalized for transfer and selectively reused based on empirical evidence. We evaluate primarily on biomedical scientific workflows, where tasks involve multi-step data analysis, domain-specific procedures, and executable evaluation. On a 39-task BioDSBench panel, we achieve 100.0% accuracy, compared with 76.9% without skills and 79.5% with SkillOpt. On BioMNIBench, AutoSkills improves the mean reward from 0.790 to 0.917, and on a larger 159-task BioDSBench evaluation reaches 94.97% accuracy versus 77.36% without skills. Beyond the primary biomedical setting, skill transfer improves performance on LiveMathematician from 0.57 to 0.68 and on OfficeQA from 0.33 to 0.60.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.