acceptodds
Under review as a conference paper at ICLR 2027

MedSkill-E²: Evaluating and Evolving Skills for Reliable Execution in Medical Agents

Abstract

Large language model (LLM) agents increasingly use skills to access domain knowledge and complete real-world tasks. We observe that even when a skill specifies the rules and procedures needed to complete a task, agents may still fail to apply them effectively during autonomous execution, a phenomenon we call skill Enable But Underuse (EBU). This issue is particularly consequential in medical tasks, where overlooking even a single critical constraint can lead to a plausible but incorrect result. To systematically evaluate and address this phenomenon, we introduce MedSkill-E², a framework for evaluating and evolving skills for medical agents. We construct EBU-Bench, a challenging benchmark covering 12 medically relevant task groups and their associated skills, to evaluate agents' ability to use available skills autonomously. We quantify the EBU gap as the difference in skill-use accuracy between autonomous use and explicitly guided reading and application of the same skill. On EBU-Bench, all evaluated model–harness combinations exhibit EBU, with a mean gap of 43.94 percentage points (pp). For the same model, skill-use accuracy under autonomous use differs by up to 38.96 pp across harnesses. Motivated by these observations, we propose EBU-Evolve, which uses execution feedback from autonomous and explicitly guided skill use to distill effective guidance into reusable skill instructions. Relative to their pre-evolution versions, the revised skills improve skill-use accuracy during autonomous execution by 33.33–55.17 pp and continue to yield improvements when transferred across multiple models and harnesses. Our results show that skill-use accuracy under autonomous execution varies substantially across harnesses even when the model and skill are held fixed, and model rankings also change; EBU can persist even after a skill is opened. EBU-Evolve distills effective guidance into reusable skill instructions, improving skill-use accuracy under autonomous execution on unseen tasks within the same skill, and still yields improvements when the evolved skills are transferred to other models and harnesses.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.