WorkFit Bench: Benchmarking Model–Scenario Fit across Real-World Agent Workflows
Abstract
Deploying agents for real-world work requires balancing task completion against model invocation cost. Yet existing benchmarks emphasize general capabilities and aggregate rankings, offering limited guidance for choosing a model for a specific work scenario. We introduce WorkFit Bench, a benchmark of 162 executable tasks across nine categories of work, constructed from reusable Skills. We evaluate 12 models across task categories and difficulty levels, measuring completion, estimated invocation cost, token use, and elapsed time. Overall completion rates range from 19.8% to 43.8%, while estimated costs for one pass through all 162 tasks range from USD 8.83 to USD 536.54 under an idealized cache-reuse assumption. Model rankings vary across scenarios and difficulty levels, and several lower-cost models achieve completion rates close to the highest overall score. WorkFit Bench provides evidence for choosing lower-cost models when they meet a scenario’s task-completion requirements. The repository is available at https://github.com/sbench-maker/WorkFit-Bench.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.