Can Agents Design Libraries for Other Agents?
Abstract
Agents increasingly build on code written by other agents, and they reimplement rather than reuse, growing the codebases later agents must work in. To measure how well agents design libraries for other agents, we introduce LibraryDesignBench, a two-phase benchmark in which an agent implements a full-featured library from a specification that defines required capabilities and potential use cases without prescribing the design. We evaluate the library through the correctness and simplicity of programs written by three user agents from different model families. Across 242 expert-validated programming problems spanning 15 library-design tasks in four languages, agent designers reproduce the abstractions of their human-authored counterparts on eleven of fifteen tasks, and downstream agents adopt both kinds of library but underuse them, reimplementing capabilities the library already provides. Our audit identifies rigid or hard-to-use interfaces as the leading obstacles in the sampled excess-code failures. We experiment with a targeted intervention: giving designers more prescriptive agent-first guidance and having them test their library with subagents improves downstream scores and yields simpler programs. LibraryDesignBench provides both a testbed for evaluating library-design practices for agent users and an initial design baseline that improves downstream reuse.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.