DatasetResearch: Benchmarking LLM Agents For Demand-Driven Dataset Discovery
Abstract
As high-quality public data becomes scarce, sustaining model progress increasingly requires agents to discover datasets tailored to specific needs. Yet it remains unclear whether autonomous agents can reliably search for, transform, or synthesize useful training data. We introduce DatasetResearch, comprising 208 reference- grounded demand specifications, each paired with a held-out reference dataset and detailed metadata. Our evaluation combines metadata alignment with downstream task performance under few-shot and fine-tuning protocols. Across open and proprietary systems, synthesis substantially outperforms search on Reasoning- intensive fine-tuning (75.46% vs. 34.52%), whereas Knowledge-intensive fine- tuning is close (43.67% vs. 45.29%) and the search advantage is clearer in few- shot evaluation. On the difficult 20-task DatasetResearch-Pro subset, a Retrieve–Transform–Synthesize workflow reaches 32.4% normalized fine-tune performance, compared with 24.1% for the strongest pure workflow; demand- only human reconstruction reaches 52.3%, demonstrating measurable headroom. These results establish DatasetResearch as a discriminative benchmark for data-research agents while clarifying the scope and attainability of its reference- normalized scores.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.