Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Abstract
Real-world enterprise data science and analytics workflows typically require reasoning across dozens of tables, performing statistical analysis and intermediate computations, and taking actions based on findings. Established benchmarks do not represent realistic, full-scale enterprise data warehouses due to the highly sensitive and confidential nature of such data. We introduce Argo-Bench, an evaluation framework comprising 210 complex data science and analytics tasks in a simulated environment. We simulate a terabyte-scale consumer food delivery platform in New York City with grounded economics and well-known patterns ranging from fraud scenarios to marketplace financial incentives based on public data sources, peer-reviewed industry literature, and regulatory data disclosures and filings. We then export this world to an ERP warehouse modeled on the Oracle E-Business Suite schema with 235 tables and 7.49 billion rows. By withholding the simulator's ground-truth state from the warehouse projection exposed to the evaluated agent, we can design realistic tasks that require reconstructing facts through the warehouse's idiosyncrasies before acting on them. Every task has an executable reference solution that demonstrates solvability using only the warehouse projection. Argo-Bench goes beyond text-to-SQL, requiring the agent to file actions such as banning users or discontinuing promotions. The strongest of 12 frontier and open-weight models reaches a score of at least 95 out of 100 on only 32.9% of tasks and averages 59.3 points. We propose that progress on Argo-Bench demonstrates meaningful progress towards understanding, navigating, and taking action within real-world data environments.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.