FinREALBench: The Execution Reliability Gap Beyond Reasoning and Acting
Abstract
Professional agents must execute workflows correctly and deliver results supported by the evidence they use. We introduce FinREALBench, a benchmark of 100 executable financial workflows spanning ten professional categories. Rather than assessing trajectories and answers in isolation, its evaluation contract follows evidence from source acquisition through transformation and repair into the submitted artifact. Process, Outcome, and Integrity assess distinct parts of this contract, while Oracle replay and targeted self-tests exercise the implemented evaluators. Across 16 model–harness systems, we examine where execution goes wrong and whether subsequent repairs reach the delivered answer. In a source-audited ownership task, all 16 submissions are schema-valid and arithmetically consistent, yet three retain an incorrect denominator. Same-model trajectories reveal two repair paths: correction from an observed filing passage and substitution of an unsupported value. Beyond this task, 133 of 868 schema-valid, fully complete artifacts receive evidence-consistency scores at most 0.5, spanning 56 tasks and 12 settings. These analyses characterize the Execution Reliability Gap: agents struggle to perform the work correctly, and complete submissions can remain unsupported. FinREALBench localizes failures along the evidence-to-delivery chain, identifying stronger execution skills, requirement-aligned checks, and evidence-grounded repair as complementary development targets.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.