acceptodds
Under review as a conference paper at ICLR 2027

Looking Beyond Accuracy: Surfacing Silent Failures in Agentic Data Science

Abstract

Large language models have significantly lowered the barrier to writing code. Naturally, this has also led to the development of agentic data science frameworks that promise automating the design and execution of end-to-end machine learning pipelines, from task formulation to model development. Existing evaluations of agentic data science (agentic DS) assess the downstream predictive performance and code correctness, providing preliminary insight in agentic DS capabilities. We argue that these evaluations overlook the fundamental requirements of data science: methodological data science knowledge and domain understanding. As such, only assessing downstream performance could mask methodological failures such as data leakage and inappropriate validation, limiting the development of robust agentic DS systems. In this paper, we evaluate agentic DS beyond accuracy from three different angles: 1) a large-scale analysis of failure modes in common agentic DS frameworks and models across 220 data science tasks, 2) a reusable and extensible test suite with curated data science tasks that test for six specific failures in a controlled manner, and 3) a qualitative assessment of agentic DS pipelines against expert solutions on Kaggle competitions. We find that the majority of agentic solutions contain at least one failure mode despite the complexity of frameworks. Our qualitative analysis shows that DS agent pipelines are generally lacking problem solving creativity and domain knowledge, compared to competition-winning data scientists. Moreover, we observe a positive correlation between the number of data leakage failures and predictive performance, illustrating the necessity to evaluate agents beyond accuracy. We believe that principled failure mode analysis, as presented here, should be an integral part of agentic DS evaluations, and that the insights we surface are key in developing trustworthy agentic DS pipelines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.