HandsOffBench: Can Agents Autonomously Complete Everyday Projects from Start to Finish?
Abstract
Large language models are increasingly deployed as general-purpose assistants to fulfill everyday user requests across applications and services. As their performance on individual tasks improves, a central question is whether agents can reliably complete entire everyday projects without further human intervention. We introduce \ours, a benchmark comprising 32 everyday projects in a dynamic simulated world featuring 70 web applications, 195 stateful services, and code execution environments. Collected from individuals with diverse backgrounds, these projects are adapted into executable tasks with carefully designed dependencies grounded in their practical requirements. Later stages build on earlier decisions, computed values, and generated artifacts, forming extended workflows with end-to-end runs spanning thousands of steps. We combine state checkers with rubric-based verifiers to assess subtask outcomes from the final environment state and generated artifacts, using an average of 105 scoring criteria per project. Shared criteria support comparisons of the same subtasks in isolated and end-to-end execution. We evaluate frontier models with a minimal scaffold and with widely used agent harnesses. The strongest evaluated model completes only 20% of projects with the minimal scaffold and 30% with harness support. On the subset evaluated in both settings, agents perform substantially worse on the same subtasks within complete projects than in isolation. These findings highlight the gap between subtask competence and reliable end-to-end completion of everyday projects, positioning \ours as a testbed for advancing model capabilities and agent harness design.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.