swespoke: Declare What You Want, Get a Bespoke Software Engineering Benchmark
Abstract
Coding benchmarks must rank agents, expose where they fail, and tell developers which agent to trust with their tasks. Early benchmarks were fixed task sets that humans defined, and current benchmarks mine tasks from an existing repository. Mining is slow and expensive, and at the end, users (software-engineering teams) get every task the repository yields. This process never asks what work users care about or which properties they want in tasks. Users, not benchmark authors, must be able to declare task properties. These declarations let users test the capability they care about, raise difficulty as agents improve, build an RL curriculum from easy to hard, and pay only for tasks that have the declared properties. These needs differ across users and change over time. swespoke lets users declare the properties they want in tasks, and a query optimizer turns the declaration into an optimized plan that builds the benchmark to reduce execution time and cost. We call the process task tailoring. Unlike mining, tailoring checks the declaration before it builds a task, so selection and creation run as one step and users pay for no task they throw away. Each task comes with a runnable environment, a build, and tests, and swespoke verifies the task by running its tests. We also build NDI, a navigation difficulty index, as an example of a user-defined property. NDI estimates how hard it is to find the code a task must change. Users compose benchmarks at a chosen NDI level. Compared with the baselines, swespoke speeds up task creation by 7.27 to 12.48×, reduces LLM expense by 2.39 to 4.85×, and accelerates measurement collection by 14.43 to 19.94×. The average coding-agent pass rate falls from 68% overall to 46% on the hardest 10% of tasks by NDI.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.