acceptodds
Under review as a conference paper at ICLR 2027

DeliverBench: Evaluating Multimodal Agents on Program-Verifiable Deliverables

Abstract

Much of the work delegated to multimodal agents turns a batch of heterogeneous material into a structured artifact that downstream software consumes, where a single mishandled item can make the artifact unusable. We introduce DeliverBench, 129 tasks that each supply a batch of authentic material, including phone screenshots, photographed invoices and nameplates, spreadsheets, audio, and video, and require such a deliverable. Deliverables are verified item by item against 411 rubric criteria under a fixed set of 37 general tools, with every call logged. Across 27 frontier models, the strongest agent completes only 41.1% of tasks, while human annotators scored by the same judge reach 87.0%; 53 tasks are solved by no model. The main empirical reversal is a batch–horizon inversion. Batch-processing tasks succeed at 16.0%, while long-horizon interactive tasks, commonly treated as the frontier difficulty, succeed at 46.0%. We hypothesize that interaction helps because it exposes verifiable intermediate states and clear stopping conditions; batch work offers neither, so one inconsistent item can silently invalidate the artifact. Two further findings concern how agents are measured. Countable efficiency summaries, such as tool calls, tokens, steps, and error rate, show no significant model-level association with success and should not be read as proxies for capability; a rubric-based rating of how each task was attempted does track success (r = +0.677). Tool-use strategy varies 19.6-fold and clusters strongly by vendor, yet is not significantly associated with success. Six shortcut screens and a paired cross-harness control support the validity of these measurements. We release the task pack, the rubric criteria, and the evaluation logs needed to reproduce every reported number.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.