acceptodds
Under review as a conference paper at ICLR 2027

DataResearchBench: Benchmarking Visual Data Sourcing with Generalist Agents on the Open Web

Abstract

Current agents are increasingly capable of retrieving information and synthesizing answers, but many research workflows require constructing data resources that are not already available in a suitable form. We introduce DataResearchBench to evaluate agents' capability of transforming high-level requirements into structured visual data collections through source discovery, acquisition, verification, and annotation. The benchmark contains 46 tasks across seven domains and requires agents to acquire images from public web resources without a predefined candidate pool. It also studies reference images as visual specifications for requirements that are difficult to express through text alone. We evaluate collected data using duplicate and provenance checks together with human-reviewed task-specific rubrics that separate base validity from additional quality. Experiments with five generalist agents (e.g., Codex, Claude Code) reveal a substantial gap between finding valid examples and constructing consistently valid collections. The strongest agent reaches 78.3% BPR (Base Pass Rate) under visual specification, yet only 45.7% of tasks contain five valid images. Visual specifications improve most agents, while fixed-backbone comparisons reveal substantial framework-dependent variation in collection quality and sourcing behavior. Larger collection requests yield more valid images overall but lower validity and quality at the largest scale, and multi-task execution further exposes differences in agent orchestration. DataResearchBench provides a testbed for studying how specification, framework, and workload shape reliable visual data construction. We will open-source the dataset and code and continuously update DataResearchBench to support emerging research needs across a broader range of data types.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.