acceptodds
Under review as a conference paper at ICLR 2027

Assess, Answer, Audit: A Point-in-Time Benchmark for Evidence-Grounded Financial Agents From Earnings Calls

Abstract

When assisting financial analysts, agents may engage in multi-turn interactions as analysts clarify and refine their requests. In this scenario, agents are required to interpret evolving requests, make decisions based on available evidence, and evaluate the reliability of answers. Yet existing benchmarks rarely assess these capabilities jointly in real multi-turn financial interactions. Inspired by earningscall dialogue, we introduce AAA-FIN (Assess–Answer–Audit for Finance), a benchmark built on real analyst–management exchanges, preceding Q&A, and the latest pre-call SEC 10-K. Three linked tasks, ASSESS, ANSWER, and AUDIT, evaluate question interpretation and evidence-based response decisions, complete and grounded answer generation, and judgments of historical answers’ evidence support and responsiveness. Together with complementary metrics, these tasks provide a multidimensional profile of financial agents’ strengths and weaknesses. We develop a reusable annotation pipeline combining LLM proposals, critic feedback, and deterministic validation, followed by review and adjudication by three domain experts. Evaluation of five API-accessible and open-weight models reveals uneven strengths across tasks, with no model consistently leading across the evaluated capabilities. Persistent gaps in evidence assessment, answer completeness, and citation coverage identify priorities for improving financial agents.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.