Agents-Eval: A Lifecycle Platform for End-to-End Agent Evaluation
Abstract
Language-model agents are increasingly used for complex tasks in software engineering, information seeking, and scientific research. Evaluating these agents is essential for reliable comparison and identifying where further improvement is needed. Yet comparing agents requires aligning benchmark protocols, agent harnesses, and execution environments. Researchers often repeat this benchmark-specific integration and configuration, incurring substantial engineering effort while leaving evaluation conditions inconsistent across studies. Existing evaluations also rely largely on aggregate scores and broad task categories, offering limited insight into which capabilities tasks require and at what demand levels. We introduce Agents-Eval, an open-source platform that addresses these two challenges. First, the platform provides reusable benchmark integrations and a unified evaluation workflow, standardizing execution conditions and lifecycle management while preserving benchmark-native interaction and scoring semantics. This allows researchers to reuse maintained evaluation setups across models and harnesses. Second, it introduces a capability translation matrix that decomposes heterogeneous tasks into six shared capability dimensions and assigns an explicit capability demand level to each. These annotations describe task requirements independently of agent responses and support comparisons across demand levels, models, and harnesses. Our experiments reveal shared capability demands across benchmark domains, harness gains that vary with model choice and task requirements, and recurring environment, service, and verification failures. Agents-Eval combines reusable evaluation workflows with fine-grained capability analysis to guide improvements in models, harnesses, and evaluation infrastructure.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.