Learning a Behavioral Skill Basis for AI Agents
Abstract
Populations of AI agents are usually summarized either by aggregate benchmark scores, which say nothing about an individual task, or by latent embeddings whose coordinates carry no fixed meaning. Neither gives a reusable representation of what a task requires and what an agent can do. We learn a compact behavioral skill basis directly from item-level outcomes, using binary success records and frozen task embeddings alone, with no skill taxonomy, no rationales, and no per-task parameters. Each task is a sparse nonnegative combination of demands on the basis, each agent carries abilities on the same coordinates, and the predicted log-odds of success is the agent's global ability plus a demand-weighted gap between ability and difficulty on each demanded skill. We then ask whether the coordinates behave as measurements, first showing that a task set determines only the profile differences its demands vary, so a task set that covers every skill equally can still leave a relative skill difference unmeasurable. Passing every prediction through 8 to 32 shared coordinates makes it exactly decomposable by skill and costs at most 0.3 accuracy points against dense amortized item response theory on EmbedLLM and RouterBench. On a controlled pool of agents, the learned profiles align strongly with withheld metadata about domain-specific training differences, which were never used during learning, reaching a selectivity AUC of over six retrainings against a shuffled null of . An unseen agent is placed in the frozen basis by fitting one number per skill, reaching accuracy from responses with all-mpnet-base-v2, against when it is included in training. Refitting an agent before and after further training on a known mixture moves its profile toward that mixture in of adaptations. Agent populations therefore contain reusable low-dimensional behavioral structure that can be learned from outcomes and used to measure how agents differ and how they change.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.