acceptodds
Under review as a conference paper at ICLR 2027

ITERACT-BENCH: Benchmarking Agents for Interactive Data Analysis Workflows

Abstract

LLM-based data agents increasingly support analytical workflows, yet existing benchmarks typically provide requirements upfront or evaluate interaction without testing whether the resulting decisions remain consistent across dependent analyses. In practice, agents must acquire analytical requirements through clarification and feedback, then apply them as persistent constraints on subsequent computations. We introduce IterAct-Bench, a benchmark for evaluating these capabilities in interactive data analysis. Built from real-world Kaggle workflows, IterAct-Bench contains 86 tasks comprising 554 analytical stages over 727 data files. A user simulator provides specification-guided clarification and correction, while independent executable verifiers assess stage-level correctness without exposing gold outputs to the simulator. We evaluate data agents using seven frontier LLMs and four agent harnesses. The strongest configuration achieves an Average Pass Rate (APR) of 70.9%, leaving substantial room for improvement in analytical reliability. Higher APR is strongly associated with proactive clarification. Agents that seek missing requirements before submitting results tend to perform better, and this behavior varies substantially across models and harnesses. Performance also generally declines as workflows contain more stages and deeper dependencies, highlighting the difficulty of maintaining analytical correctness across complex workflows. IterAct-Bench enables systematic evaluation of how agents acquire requirements through interaction and preserve them as constraints on downstream analysis. The data and code are available at [https://huggingface.co/datasets/blue-orbit-472/IteractBench](https://huggingface.co/datasets/blue-orbit-472/IteractBench).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.