acceptodds
Under review as a conference paper at ICLR 2027

WorkFit Bench: Benchmarking Model–Scenario Fit across Real-World Agent Workflows

Abstract

Deploying agents for real-world work requires balancing task completion against model invocation cost. Yet existing benchmarks emphasize general capabilities and aggregate rankings, offering limited guidance for choosing a model for a specific work scenario. We introduce WorkFit Bench, a benchmark of 162 executable tasks across nine categories of work, constructed from reusable Skills. We evaluate 12 models across task categories and difficulty levels, measuring completion, estimated invocation cost, token use, and elapsed time. Overall completion rates range from 19.8% to 43.8%, while estimated costs for one pass through all 162 tasks range from USD 8.83 to USD 536.54 under an idealized cache-reuse assumption. Model rankings vary across scenarios and difficulty levels, and several lower-cost models achieve completion rates close to the highest overall score. WorkFit Bench provides evidence for choosing lower-cost models when they meet a scenario’s task-completion requirements. The repository is available at https://github.com/sbench-maker/WorkFit-Bench.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.