acceptodds
Under review as a conference paper at ICLR 2027

WorkThrough-Bench: Benchmarking Agents on Long-Horizon Productivity Work

Abstract

Long-horizon productivity work is becoming the proving ground for AI agents. Evaluating agents on such work remains challenging, as existing benchmarks often focus on minute-scale tasks, narrow work settings, or one-shot generation. We present WorkThrough-Bench, 86 tasks from professional requests across dozens of industries, spanning deep research, deliverable creation, and application building. Multi-artifact deliverables are independently graded through rule-based checks, LLM judgments, and behavioral probes. Across six frontier models and two agent harnesses, execution averages 122.8 minutes per run. Agents satisfy an average of 68.0% of the evaluation criteria per task, yet task pass rates remain below 11% in every model–harness group, and 54 of the 86 tasks are never completed in full by any group. Examining this gap between partial progress and reliable delivery, trajectory analysis reveals self-contented completion, where agents conclude that a task is complete based on checks of their own implementation and remembered requirements, without verifying the original task contract. We introduce feedback that redirects verification toward the original requirements and prompts corrections to overlooked discrepancies. Using fixed feedback, we improve the mean task pass rate by 26.7% relative to the no-feedback baseline, with trajectories showing renewed inspection and targeted repairs. These findings highlight the potential of further feedback design to improve agents' completion judgments and final delivery.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.