acceptodds
Under review as a conference paper at ICLR 2027

DataClawBench: An Agent Benchmark for Exploratory Real-World Financial Data Analysis

Abstract

Realistic analysis often begins without knowing where the relevant evidence lies, raising a fundamental question for autonomous data analysis: can agents discover and use the required evidence with limited human guidance? However, existing benchmarks typically evaluate such agents in human-guided settings, providing selected data sources, explicit data schemas, or cleaned data, thereby understating the exploratory burden. To evaluate autonomous exploratory analysis in a more realistic setting, we introduce DataClawBench, a benchmark built from financial think tank consulting scenarios where agents must independently explore unfamiliar, noisy, cross-domain data and produce verifiable conclusions. DataClawBench provides a unified real-world data environment containing approximately 2.06 million records across enterprise, industry, and policy domains while preserving native data imperfections. On this environment, it defines 492 multi-step cross-domain tasks that withhold data source hints, complete schema documentation, and explicit noise descriptions. To reveal exploratory progress and failure modes beyond final outcomes, we annotate critical intermediate milestones and introduce a process-oriented evaluation framework. A systematic evaluation of ten advanced LLMs under the OpenClaw agent reveals that exploratory data analysis breaks agent reliability: more exploration does not reliably translate into task-relevant progress or correct final answers

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.