SWE-Skills-Bench: Do Agent Skills Actually Help in Real-World Software Engineering?
Abstract
Agent skills, structured procedural knowledge packages injected at inference time, are increasingly used to augment LLM agents on software engineering tasks. However, their real utility in end-to-end development settings remains unclear, and no standard procedure tests an individual skill before installation. We present SWE-Skills-Bench, a skill-centric, requirement-driven benchmark that measures the extent to which individual agent skills improve performance on real-world software engineering (SWE) tasks. We start from 98 public SWE skills, each paired with a real GitHub repository pinned at a fixed commit. For each skill, our pipeline (i) reverse-engineers ten requirements specific to that repository, (ii) turns each requirement's acceptance criteria into deterministic, execution-based verifiers, and (iii) builds a reproducible container environment. Under each of four agent–model configurations, we run every task once without the skill and once with it, yielding 3,920 paired evaluations. For each requirement, we compute a paired delta, i.e., the change in outcome when the skill is added. We then aggregate these deltas across a skill's ten requirements into a per-skill utility profile, and use a statistical test to classify each skill as Golden (significantly helpful), Ineffective (no significant effect), or Harmful (significantly detrimental). Our results show that skills provide limited benefits: 92 of 98 skills have no statistically significant effect, and adding skills raises the average pass ratio only from to . The few helpful skills supply procedures that the agent would otherwise omit (gains of up to points). In contrast, harmful skills anchor the agent on their own templates and override the task's explicit instructions (losses of up to points). Moreover, a skill's utility depends on the agent harness and model: 40 of 98 skills flip the sign of their effect between two harnesses running the same model. These findings suggest that skill utility is an empirical, configuration-dependent property rather than an intrinsic one, and SWE-Skills-Bench provides a testbed for testing skills before deployment. Our code is available at https://anonymous.4open.science/r/SWE-Skills-Bench.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.