From Solving Data Tasks to Learning a Data Space: Benchmarking Learning Capabilities of Data Agents
Abstract
Data agents are increasingly deployed in enterprise analytical platforms, where they serve sustained workloads of queries over the data space. These workloads often contain queries that share procedures, depend on domain knowledge specific to the data space, and are affected by changes in it. Without learning capabilities, data agents may repeat costly exploration for each query, miss knowledge they have already found, and produce wrong answers from outdated experience. Existing data agent benchmarks evaluate each query independently and do not measure the learning capabilities of data agents. In this work, we introduce DataIntern, a benchmark that evaluates the learning capabilities of data agents through: (1) 15 real-world enterprise-level data spaces and 742 queries organized into 84 workloads that cover three complementary learning scenarios: procedural reuse, data space understanding, and data space changes; (2) a feedback-aware agentic construction framework, in which real enterprise feedback guides agents to build data spaces from public data and to verify by execution that each query needs this domain knowledge; (3) a controlled evaluation protocol that evaluates the learning capabilities of data agents under different experience conditions and clearly reveals how much they gain from experience and at what cost; (4) extensive experiments across thirteen frontier LLMs, six agent harnesses, and four memory mechanisms that reveal key insights for building future data agents with stronger learning capabilities. Our results show that reusing procedures cuts tool calls by 81 % and cost by 60 % on DataIntern-Full. Compared with unrelated experience from the same data space, experience that contains this domain knowledge raises accuracy by 15–28 points for twelve of thirteen models on DataIntern-Lite. Data space changes remain the hardest scenario. Few data agents answer correctly before and after a change, and most failures come from applying knowledge where it does not hold. Code and data are available at https://anonymous.4open.science/r/DataIntern-E53C.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.