acceptodds
Under review as a conference paper at ICLR 2027

Aggregate Scores Mask Generalist Agent’s Skill Usefulness

Abstract

Recent advances in AI agents are moving toward generalist agents that acquire new capabilities through modular skills—reusable capabilities, knowledge, and tools that let agents perform new tasks. This raises a key evaluation question: when an agent already has many skills, what does adding another skill actually change? Standard with- or without-skill benchmarks evaluate skills in isolation, measuring whether they help on target tasks while overlooking whether they duplicate existing skills or degrade capabilities the agent already has. We reframe skill evaluation as an incremental, system-level attribution problem: we compare the agent before and after a skill addition across a grid of system state and test case context, and introduce a capability correctness (CC) metric that scores the capability atomic operations inside a skill, crediting exact and functionally equivalent invocations. Applied to a generalist agent extended with seven open-source domains as skills (320 released test cases), we find that a single skill addition produces both gains on its target tasks and regressions on previously supported ones—including redundancy and cross-skill interference that aggregate benchmarks hide. Because CC localizes these effects to specific capabilities, it also yields an actionable diagnostic signal for skill refinement, which we probe with a simple targeted pruning method. Therefore, we treat aggregate scores as insufficient to attribute the true usefulness of a skill.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.