BLUE: Benchmarks, Let Us Evolve with Agent
Abstract
As LLM agents acquire new skills, tools, and other actions, their capability space continues to expand. However, existing benchmarks are either static or generate more complex tasks only within a previously defined capability space, creating a growing mismatch between what agents can do and what benchmarks can evaluate. To address this gap, we introduce BLUE, an evolvable benchmark generator for agent evaluation. For each newly introduced capability, BLUE requires only a simple user-supplied atomic task and its corresponding evaluation. It groups atomic tasks into atomic scenarios, constructs a scenario graph to generate complex tasks, and simultaneously composes atomic evaluations into evaluators for these complex tasks. BLUE combines rule-based metrics with LLM judges and supports both single-turn and multi-turn evaluation. We construct BLUE-Bench, which covers diverse multi-capability scenarios, to systematically evaluate five representative coding agents paired with three foundation models. We further conduct a detailed analysis of BLUE key components to validate the effectiveness of the framework. All code and data are available at https://anonymous.4open.science/r/BLUE-ICLR.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.