acceptodds
Under review as a conference paper at ICLR 2027

DAVBench: Benchmarking Agentic Verification of Real-World Data Work

Abstract

Data work addresses user-specified data needs, spanning data engineering (DE), which builds and transforms datasets and data pipelines, and data analysis (DA), which derives metrics and analytical results. Reliable verification is critical for preventing faulty transformations and flawed analytical conclusions before candidate data work is delivered. Yet existing benchmarks largely evaluate agents' ability to complete data tasks, while overlooking their ability to verify whether candidate data work satisfies user requirements and is free of defects. We introduce **DAVBench** (Data Agentic Verification Benchmark), comprising **216 repository-level verification tasks of data work** with reproduced real-world defects across data engineering pipelines, analytical workflows, and delivered artifacts. Tasks involve connecting evidence across data, code, and domain documentation, tracing defects through multi-stage workflows, and distinguishing faulty from correct candidates. We evaluate verification along two dimensions: ***(1) verification quality***, which measures the accuracy and completeness of verification reports using location matching and evidence-grounded semantic judging; and ***(2) repair utility***, which measures how effectively these reports help a downstream model correct faulty candidate work under behavioral tests. Our empirical results show that current agents still struggle to verify real-world data work reliably. Even the flagship model GPT-6 Astra achieves a finding F1 of only **49.82%** and complete defect coverage on just **39.61%** of faulty tasks. Further analysis reveals that verification quality is a key bottleneck for correcting data work. When given reference verification reports, a downstream repair model successfully repairs **69.48%** of tasks, compared with **24.03%** without reports and only **37.66%** with reports from the best-performing verifier. Overall, DAVBench provides a testbed for improving the verification of real-world data work.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.