FinEvo-Bench: A Longitudinal Benchmark for Self-Evolving Agents in Professional Financial Workflows
Abstract
Agents used over time encounter recurring professional work: each case requires different evidence and judgment, while the underlying workflow can be reused. Benchmarks built from independent tasks cannot reveal whether an agent turns earlier experience into better procedures for later cases. We introduce FinEvo-Bench, a longitudinal benchmark designed around this structure. It contains 120 open-ended tasks drawn from real cases across 20 business scenes in six financial domains. Each scene contains six substantively different cases that share a professional workflow and an expert-authored rubric for task quality and financial compliance. Constructing and validating the benchmark required approximately 1,200 person-hours. Finance provides a natural test bed because recurring analyses apply shared professional and compliance requirements to heterogeneous inputs, producing case-specific analyses and conclusions. We evaluate four self-evolving agent scaffolds with Qwen3.7-Max on three independently shuffled, globally interleaved task streams. A Claude Code rubric judge backed by Claude Opus 4.6 evaluates all outputs, and paired state-reset controls estimate each scaffold’s gain from retained experience. Evolving runs score 9.33–19.37 points higher and trigger 0.12–0.44 fewer compliance issues per task than their paired controls. Paired score gains at within-scene ranks 4–6 exceed those at ranks 1–3 by 6.10–8.70 points. FinEvo-Bench measures whether retained experience improves later professional work under continued use. The dataset is available at https://anonymous.4open.science/r/finevobench-710A/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.