Kolab: Benchmarking Human-Agent Collaboration in Long-Horizon Workflows
Abstract
When humans collaborate with AI agents on long-horizon tasks, what happens to the agents' contributions? For solving long-horizon tasks, humans place greater importance on ongoing interaction with AI agents (e.g., revising plans) than on an agent's ability to complete the task in a single attempt. However, most agentic benchmarks have focused more on automation. Strong performance on these benchmarks does not mean that an agent can collaborate effectively when deployed in the real world. Professional workflows offer a useful setting for examining this question, as they often involve interdependent steps. To address this gap, we introduce Kolab, a new dataset of 100+ hours of screen recordings of 45 professionals from 36 occupations collaborating with AI agents on realistic, long-horizon problems. The dataset also contains collaboration and output quality judgments from the participating experts. We evaluate (1) how well agents collaborate with professionals and (2) whether they can serve as verifiers of human–agent collaboration. With Kolab, we find that frontier models, both agentic and non-agentic, tend to overestimate collaboration with experts and struggle to distinguish between high- and low-quality collaboration. As frontier agents cannot yet understand long-horizon collaboration with experts, we position Kolab as a valuable dataset for systematically studying human–agent collaboration grounded in real professional workflows.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.