acceptodds
Under review as a conference paper at ICLR 2027

QuantitativeFinance-Bench: Benchmarking AI Agents on Real-World Quantitative Finance Tasks

Abstract

Evaluating language models on quantitative finance tasks is challenging because such tasks require simultaneous mathematical rigor, domain expertise, and data engineering in a single executable artifact. We introduce QuantitativeFinance-Bench (QF-Bench), an execution-grounded benchmark for the computational core of professional quantitative finance. We have four main contributions. First, we curate 86 practitioner-authored tasks spanning derivatives pricing, fixed income, credit, risk management, factor research, backtesting, microstructure, and financial NLP. Each task is executed in a network-isolated Docker container with real financial data. Second, we introduce a hierarchical verification framework with more than 3,300 assertions across 86 tasks. These assertions encode domain-specific checks, including no-arbitrage bounds, put-call parity, and convergence diagnostics that are not captured by generic code-level tests. Third, we define Finance-Zero, a non-agentic baseline for difficulty calibration, and introduce automated error attribution for Computation, Convention, and Mislabeling errors. Fourth, we evaluate 38 Finance-Zero base-model configurations and 9 production CLI-agent configurations, finding that the best system achieves 61.7% pass@1 while the single-call baseline reaches 42.6%. A run passes only when all required deliverables and verifier checkpoints succeed. The benchmark and evaluation resources are available at https://anonymous.4open.science/r/QuantitativeFinance-Bench-07BB.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.